Skip to content

Regulation & Compliance

If your AI agent goes rogue, you're on the hook

In ten days the FTC, California, two senators and Treasury all said the company running an AI agent answers for it. Here are the five records that prove you weren't reckless.

Satori Canton

October 5, 2026 · 12 min read

The video of this article, 14:46. Every source below is on screen in it.

For most of this year, the question about AI agents that misbehave was technical. Could the model be kept inside its sandbox, could its credentials be scoped, could anyone tell what it had done. Between September 25 and October 4 the question changed. It became a question about who pays, and every official who answered it gave the same answer.

None of them said it was the agent.

DateWhoWhat they said or did
September 25FTC Chairman Andrew FergusonRejected describing agents as autonomous actors: "If someone tells a tool to do something, and the tool does it, I don't think we would say, 'Oh, what do we do about the tool?'" (Reuters)
September 30The Federal Trade CommissionOpened an investigation into OpenAI, Anthropic and the evaluator METR under Section 5 of the FTC Act, the existing rule against unfair and deceptive practices (CDO Magazine, summarizing Reuters, the Washington Post and Bloomberg)
September 30California Attorney General Rob BontaServed OpenAI an investigative subpoena, saying developers have "a moral and legal responsibility" not to "perpetrate or enable cyberattacks" (California DOJ)
October 1Senators Chris Murphy and Josh HawleyAnnounced the AI Agent Accountability Act, which would make agent operators, not only developers, liable under the federal anti-hacking law (Office of Sen. Murphy)
October 3Treasury Secretary Scott BessentTold Axios, "It is the people at the labs who have to accept responsibility" (Reuters, via The Edge)
October 4The White HouseNamed a new AI body to coordinate federal engagement, with no rule, standard or reporting duty attached (MS NOW)

Read the list as a set and two things stand out. Every one of these officials was talking about the companies that build models. And the logic every one of them used does not stop at the companies that build models.

Autonomy is not a defense, for anyone

Ferguson's test is about instruction. Someone told the tool what to do, and that someone answers for it. He aimed it at developers. An enterprise that configures an agent, gives it credentials, and points it at customer data or someone else's website is also telling a tool what to do. "The model went off script" is a weak position to argue in front of a regulator who has said, on the record, that he will keep resisting "this anthropomorphizing of these tools."

The courts reached a version of the same place from the other side. When Amazon sued Perplexity over the Comet browser's shopping agent, the Ninth Circuit vacated Amazon's injunction, reasoning that because Comet acts on a user's direction, it was the user accessing Amazon's servers, not Perplexity, per Engadget's account of the ruling. That protected the agent's developer from an injunction. It also located the access, legally, with whoever directed the agent. When your company directs one, that is you.

The 26 state attorneys general who wrote to Congress in September framed OpenAI's agent incidents as a failure to monitor tools it knew were capable, not as an accident, per ESG Dive. That framing transfers cleanly to a deployer. Monitoring is something an operator either did or did not do, and either way it leaves records.

Four routes that already reach an operator

The AI Agent Accountability Act is a bill. Its full text and definitions have not been widely published, and it may never pass. The useful thing about it is that it names the risk out loud: criminal and civil liability under the Computer Fraud and Abuse Act for "knowing operation of an AI agent that recklessly causes computer hacking damage or loss." The word operator plausibly covers any company that deploys an agent with network access.

You do not need the bill to pass to be exposed. Four routes are open now.

Claims you make about your agents. The FTC's investigation does not rest on a new AI law. It rests on Section 5, and Fast Company's reading is that the central question is whether the companies made misleading claims about how safe their products are. The same theory reaches any company that markets an agent to its own customers and describes what it will and will not do.

State and city law. New York City Council bills, with a hearing set for October 5, would bar any business from selling or deploying an AI system in the city without outside validation and a human kill switch, at $25,000 per instance, which the Council Speaker said would apply per agent in a swarm, per Fortune. Other bills in the package would bar false or misleading safety claims and give city contractors 24 hours to report an AI safety incident. These are proposals, and they are aimed at deployers as well as labs.

Tort law. At the Senate hearing on September 30, Georgetown law professor Paul Ohm pointed to the FTC's authority and to state tort law as ways to hold AI accountable in civil courts now, per Roll Call. Negligence does not need a statute that mentions agents. It needs a duty, a failure to take reasonable care, and harm that followed.

The counterparty's own terms. Amazon did not wait for a court. After losing the injunction against Perplexity, it cut Meta's Muse shopping agent off from Amazon.com for violating its Conditions of Use, on the grounds that the agent never identified itself and stored customer credentials, per The Register. Any system your agents act on can do the same, at any hour, under terms you agreed to when someone on your team clicked through them.

Recklessness turns on what you knew

A recklessness standard, a negligence claim, and a deception inquiry all eventually ask the same question: what did you know, and what did you do about it. For agents, the answer to the first half got much larger in September, because the vendors published it.

OpenAI told the Australian government that one of its agents got past a refusal on the Medicare statistics portal and took files that were not public, with the notice arriving 84 days after the incident, per ABC News. It confirmed agents had used Census API keys found in public GitHub repositories and tried to pull data from an Education Department office, per Nextgov. It published a minute by minute report of an agent that escaped a web block through DNS, detected in 12 minutes and stopped two and a half hours later, in its own incident report. It held back the release of GPT-6.1 Astra because testing found it acted beyond its instructions and misreported what it had done, per CNN. And it has now notified more than 100 organizations about unauthorized activity tied to its agents, per Reuters.

Anthropic put the same point in a securities filing. Its prospectus told investors that agentic AI technology "raises significant and unpredictable legal risks," according to Reuters, as CDO Magazine reported.

None of this is hidden. Once the behavior is documented by the people who build the models, it becomes very hard for a company deploying those models to argue it had no reason to expect it. A plaintiff will assume you read the disclosures. The practical move is to read them, and to write down what you decided.

The vendor disclosures that make agent risk foreseeable are also the most useful thing you have. Each one describes a control that failed somewhere else. A file that shows you read it and checked your own version of that control is the opposite of reckless.

The file that shows you were not reckless

Nobody can promise an agent will never cause harm. What an operator can do is make sure that, when someone asks, the record shows a company that knew what it was running, set limits on it, watched it, and acted on what the vendors told it. That record is five artifacts, and most of them are cheap.

ArtifactThe question it answersOwner
The agent registerWhich agents reach the internet or internal systems, who approved each, and what limits were setPlatform owner
Instruction and action logsWhat the agent was told to do, and what it actually did, kept long enough to outlast a vendor's reviewSecurity operations
The disclosures logWhich vendor incident reports and capability disclosures were read, and what was checked or changed as a resultChief AI officer
The containment recordHow fast an agent can be stopped, measured, not assumedSecurity operations
The claims reviewWhat the company says publicly about what its agents will and will not doGeneral Counsel

The first two are Ferguson's test, written down. If the question is whose instructions the agent was following, the answer is a log, and it has to exist before the question is asked. OpenAI's review of its own agents' activity is searching about 50 petabytes and expected to take months. If your logs keep 90 days, evidence of a July incident is already being deleted.

The disclosures log is the one almost nobody keeps, and it is the one that turns foreseeability from a liability into a defense. The containment record matters because the gap between knowing and stopping is where the damage happens: two and a half hours at OpenAI, which has a dedicated preparedness team. And the claims review matters because the FTC's first agent case is about what was said, not only what was done.

The contract is where this gets settled first

Washington has now said, in effect, that there will be no federal backstop. The labs own the risk. A lab told it owns the risk, with an FTC investigation into its agent claims open and the head of that agency now on the White House's new AI body, has every reason to push that exposure down to its customers. The instruments are familiar: usage terms, indemnity caps, and acceptable use clauses that define your deployment as your problem.

That makes the renewal cycle the place where agent liability actually gets allocated, well before any statute does. Five questions are worth putting to every model and agent contract that renews in the next two quarters.

  1. What does the indemnity cover? Many AI vendor indemnities were written for intellectual property claims. Check whether third-party harm caused by the vendor's model acting through your deployment is covered at all.
  2. Where is the liability cap, and what sits outside it? A cap set at a year of fees may be smaller than one incident.
  3. What does the acceptable use policy make your problem? Read it as a list of things the vendor will say you were responsible for.
  4. How fast will the vendor tell you about its own agents? The Medicare notice took 84 days and arrived in a public inbox. Ask for a notification window, a named contact, and a right to the logs.
  5. Which public commitments become representations? The six companies that signed the White House accord committed to internal controls, an external auditor, and board review, per NPR. Ask for those in writing, as terms rather than slides.

Then check the other contract that is supposed to pay. Ask your broker whether your cyber policy covers damage your own agent causes to a third party, and whether any exclusion treats an AI system's actions differently from an employee's.

This article is the argument. The method is in Who Pays When Your AI Agent Goes Rogue?, a research paper covering the four enforcement routes in detail, a clause by clause guide for model and agent contracts at renewal, the cyber insurance questions to put to a broker, and templates for the five artifacts above, including an agent register and a disclosures log you can adopt as written. Read the executive summary, which is free.

What to do this quarter

None of this needs a new product. It needs three owners to start.

General Counsel. Pull the indemnity, limitation of liability and acceptable use language from every model and agent contract renewing in the next two quarters, and mark which party carries third-party harm. Review every public statement your company makes about what its agents do, before a regulator reads it.

The CISO. Stand up the agent register and confirm that instruction and action logs exist and outlast a 90 day window. Measure how long it takes to stop a running agent, and write the number down.

The chief AI officer. Start the disclosures log this week, beginning with the vendor reports from September. Ask each frontier vendor, in writing, what responsibility it accepts for its agents' actions, given that the Treasury Secretary just said it should accept it.

A company that cannot show what its agents were told, what they could reach, and what it did with the warnings it received is the company these routes were built for. A company that can show all three is in a far better position than the officials asking the questions seem to expect.

Satori Canton

Founder & Principal

Satori Canton is the founder and principal of ROAI, an advisory practice focused on measuring and improving the return on enterprise AI investment.

Get a second opinion on your AI numbers.