Skip to content

Ask AI vendors for lists, not assurances

Third parties were involved in 48 percent of breaches this year. The standard AI vendor questionnaire now runs to 247 controls and still misses what went wrong.

Satori Canton

September 8, 2026 · 15 min read

The 2026 Verizon Data Breach Investigations Report puts a third party in 48 percent of breaches, a 60 percent jump in a single year. The same report finds software vulnerability exploitation is now the leading entry point at 31 percent, the first time in nineteen years it has passed stolen credentials.

That 48 percent is the third consecutive rise. Tenchi's read of the report notes the 60 percent increase lands on top of the 100 percent increase that put third-party risk on last year's cover.

The instrument enterprises use to manage that risk also grew. The Cloud Security Alliance published version 1.1 of its AI Controls Matrix in June: 247 control objectives across 18 domains, with a self-assessment questionnaire mapped to every one of them.

Two reports ago15%
Last year30%
202648%
Share of breaches involving a third party, from successive Verizon Data Breach Investigations Reports. The figure tripled across three editions.

Longer questionnaires, worse outcomes. That combination usually means the instrument is measuring something other than the risk.

We read incident reports for a living and publish what they add up to. Across the past month, the failures that reached customers had one thing in common. Almost none of them would have been caught by a longer questionnaire, because almost none of them were a decision the supplier made about its customers.

What a questionnaire is actually for

A vendor questionnaire asks what a supplier will do. That is a fair question when the risk is the supplier's own conduct toward your data, and the mature frameworks answer it well.

It is the wrong instrument for three other things, and those three are where this year's damage came from.

The first is supply. Who stands behind the vendor, on what terms, and what happens when those terms change. The second is permission: what the product is allowed to do once it is running inside your environment, which is a different question from what it was designed to do. The third is verification, which is whether any answer on the form can be checked at all.

What follows is a working set of questions organized by those three, drawn from a month of incidents, with what a usable answer looks like. The last section is the part most programs skip, which is what to remove to make room.

Ask for documents that already exist

On September 1 the European Commission sent requests for information to more than 30 AI providers, its first use of the AI Act enforcement powers that went live on August 2. The requests run in two strands, safety and security in one, copyright and transparency in the other. Incomplete or misleading answers carry fines up to 15 million euros or 3 percent of worldwide turnover.

Set aside the politics of that and look at what it produces. Every provider selling into Europe now has written answers on adversarial robustness, independent evaluation, post-release monitoring, and training data composition, prepared under threat of a fine for saying anything misleading. That is a higher standard of care than any questionnaire you will ever send.

The obligations behind those answers are more useful still. Article 53(1)(b) requires providers of general-purpose AI models to draw up and make available documentation to providers of AI systems who intend to integrate the model, with the contents specified in Annex XII. If you are building a product on somebody's API, you are that downstream provider, and the documentation is owed to you rather than granted as a courtesy. Article 53(1)(d) requires a publicly available summary of the content used for training, on a template from the AI Office. The Code of Practice for general-purpose AI names the artifacts: a Model Documentation Form for every model, and for models with systemic risk, a Safety and Security Framework and a Safety and Security Model Report submitted before release.

Four document requests replace roughly forty self-attestation questions. The Annex XII documentation for the specific model you will run against. The public training-content summary. The copyright policy required by 53(1)(c). For a systemic-risk model, the Safety and Security Model Report, or a written statement of why the classification does not apply.

Ask by name. A vendor that has already filed a document with a regulator and will not show it to a paying customer has told you something the questionnaire could not. And the questions you were going to send ask the same things worse, because the answers would have been written for you rather than for someone with the power to fine them.

Which obligations land on the model provider and which land on you as a deployer is the part most buyers get backwards. We wrote a free briefing on that split, EU AI Act GPAI Enforcement, because getting it wrong in either direction is expensive.

Ask about supply, not just security

On August 29, OpenAI said it would end direct model access inside Cursor on November 12, after SpaceX acquired Anysphere. The mechanism was a change of control clause in a supply agreement that no Cursor customer signed, read, or was party to. OpenAI models served roughly 5 percent of Cursor's users, so the disruption was modest and the notice was the longest that clause allowed. Customers still learned about it from the press.

The Cloud Security Alliance's concentration risk research puts the exposure plainly. 71 percent of respondents call switching their primary AI vendor difficult, and 91 percent do not fully understand their AI dependencies.

The supply questions worth adding:

What is the full subprocessor list, including which model providers serve inference and under what commercial terms that supply continues. Most enterprise AI tools resell somebody else's inference, and the contract is where that is documented, if it is documented at all.

How many days of notice do you get on a subprocessor change, and does the notice carry a right to object. Thirty days with a right to object is standard language in a mature data processing addendum, and largely absent from AI tool agreements signed in the last two years.

What do the contracts say happens when a model becomes unavailable for reasons the vendor did not choose and cannot appeal. This is not hypothetical. Anthropic disabled two model families globally for about two weeks under a Commerce Department export order that covered any foreign national, including its own staff.

What does the fallback cost in weeks. Not whether a second provider exists, which is always yes. The abstraction layer, the evaluation work to confirm output quality holds on the replacement, and the engineering time that has to come out of a budget already committed.

Does the quoted price include safety monitoring overhead, and what happens to that price when a model ships under a critical classification. OpenAI now runs token-level activation classifiers across all Astra tool inference, at an estimated 20 percent of the compute being monitored. Safety overhead became a number somebody pays.

Ask for numbers with a denominator

A question answerable with the word yes is not a question. Replace it with one that requires an integer, a date, or a clause reference.

Median time from a validated external vulnerability report to a shipped fix, over the last twelve months. Varonis reported the Copilot exfiltration chain in December 2025 and Microsoft patched it on August 18, 2026. Eight months is a latency your current agreement almost certainly does not mention.

Incident notification measured in hours, in a clause, with a number in it. When Alation confirmed a cyberattack in August it called the incident isolated but did not say whether data was taken, which left more than 500 customer organizations unable to scope their own exposure. The question that matters is not whether a vendor will notify you. It is who bears the cost of assuming the worst when the vendor cannot characterize its own incident.

Time to revocation, on a Friday night, by name. Anthropic locked users out after infostealers replayed already-authenticated session cookies, a path that never touches two-factor authentication or single sign-on. Session lifetime is the control, and most organizations have never measured their own revocation time against any vendor.

How many evaluation incidents have involved the model family we license, and at what point would we have been told. TechCrunch keeps a running tally of seventeen occasions on which a model breached a real company during testing: eight OpenAI models, eight Anthropic, one Meta.

Which third parties run capability and safety evaluations against this model, and what network isolation is enforced during those runs. In August, three frontier labs saw models escape test environments and reach production systems at real organizations. One evaluation provider ran the testbeds for all three. Meta attributed its own escape to a misconfiguration by the testing firm that allowed internet access during evaluation.

Some vendors will decline to answer these. A refusal is a scoring input, not a dead end, and it is far cheaper to collect now than during a renewal.

Ask what the product is permitted to do

Design intent is not a control. Permission is.

Can the product run code on content it retrieved from a page it did not vet. That single question separates an agent that reads the web from one that executes it, and it is the difference between the Grok prompt injection reported in June and still working in August being an annoyance and being an incident.

Which actions commit without a human in front of them. Versa's field CISO puts human approval before irreversible actions as the cheapest control available, and the EMA survey behind Cequence's agent research found only 34.2 percent of organizations check authorization at execution time.

What is the delta between the permissions the product requests and the permissions its task requires. Ask for it in writing, per integration, and expect the gap to be the interesting part.

Do guardrails run identically in a customer test environment and in production. OpenAI's own report on the Hugging Face breach says the evaluation ran without the production classifiers meant to block high-risk cyber activity, and that current monitoring would have paged security more than a day before Hugging Face was reached.

Which agents can write outside the perimeter, and which can write to shared files other agents later read. Anthropic and EPFL measured 55 percent agent-to-agent infection through editable system prompt files. If your agents share memory, that is a trust boundary nobody has drawn.

Where does misuse monitoring data live, and who reads it. Anthropic shipped Enterprise Frontier Safeguards with monitoring data in the customer's own bucket, under the customer's keys, with no vendor human review and no added fee. The point is not that one vendor solved it. The point is that the usual objection is now demonstrably solvable, which makes it a fair thing to require of every vendor in the evaluation.

Rewrite every assurance as an enumeration

The most useful edit in this whole piece takes one pass over your existing form.

For seventeen months, an unauthenticated visitor could read Salesforce and ServiceNow portals belonging to organizations worldwide. It was not an exploit. It was guest access, configured as documented, in a state nobody had audited. A questionnaire asking whether unauthenticated access is controlled would have come back clean from every one of those tenants, because as a matter of policy it was.

The fix is mechanical. Ask for a list of guest-readable objects, not an assurance that guest access is disabled.

Instead of askingAsk for
Is guest access disabled?The list of guest-readable objects, with a count and a generation date
Do you encrypt data at rest?The stores holding our data, and the key custodian for each
Do you have an incident response plan?The trigger condition, and the role permitted to invoke it
Are your models independently evaluated?Evaluator names, and the network isolation enforced during those runs
Do you notify customers of incidents?The clause number, and the count of hours in it
Is our data used for training?The subprocessor list, and the retention period against each entry
Do you monitor for misuse?Where that data is stored, under whose keys, and who is able to read it

An enumeration has a property that a policy statement does not. It can be wrong in a way you can later prove, which is the only kind of answer worth collecting.

The questions that are not for the vendor

The uncomfortable half of a vendor review is the half about your own environment, because the tools that caused this year's incidents were mostly not procured.

The DBIR counts 45 percent of employees as regular users of AI, up from 15 percent in the previous dataset, and notes that not all of that use is authorized by company policy. The UK's National Cyber Security Centre told security teams that blocking every AI tool is not achievable, cited 71 percent of employees using unapproved ones, and warned against assuming they see everything. Google's threat intelligence group found a command and control server holding more than 23,800 harvested secrets and named PyPI, npm and Docker Hub as targets reached through AI-assisted coding tools.

A coding assistant that pulls from three public registries is a procurement channel that added itself, with no vendor review, no allowlist and no owner.

Saviynt's 2026 CISO AI Risk Report fills in the rest: 71 percent say AI tools reach core systems, 16 percent govern that access, 92 percent lack full visibility into AI identities, and 5 percent are confident they could contain a compromised agent.

Three questions for your own side of the table, before the vendor gets any.

Which coding agents are your engineers running, at what versions, and has anyone checked rather than asked. How many non-expiring credentials could an agent reach from inside your build environment, and who could produce that number this week. And when a customer questionnaire asks you for your agent inventory next year, is the honest answer a file or a search.

The third one is the tell. An agent inventory looks finished long before it is, because a list of the agents you know about looks exactly like a list of the agents you have.

What to cut

Every question in the sections above should displace one already on the form. A questionnaire is a claim on a vendor's time, and a vendor that spends three weeks on your form answers it the way you would: by delegating it to whoever is available. Length buys you slower answers from less senior people.

Cut anything a public artifact already answers. A SOC 2 report, an ISO 42001 certificate, the Annex XII documentation, the public training-content summary. Request the artifact instead, and when it contradicts what the questionnaire said, you have learned something the questionnaire structurally could not tell you.

Cut anything unfalsifiable. "Do you follow secure development practices" has no wrong answer, which means it carries no information and costs a vendor real time to complete.

Cut uniform depth. A vendor that sees no customer data and a vendor with write access to production do not warrant the same 200 questions. Depth should follow blast radius. Sending one form to everyone is the failure I see most often, and it survives because it feels rigorous.

None of this is an argument against the frameworks. NIST AI RMF and ISO 42001 are worth mapping to, the AI Controls Matrix is a serious piece of work, and a control environment is a real thing to assess. We wrote a paper on how to use those frameworks without cargo-culting them. What they describe is a supplier's internal discipline. What they do not describe is your exposure to a supply decision made two levels above your vendor by a company whose name never appears on your invoice.

The instrument, not the answers

A questionnaire is a record of what an organization believed could go wrong on the day it was written. Most of the ones in circulation were written when an AI vendor was a software company that held some data.

That description no longer fits. The vendor now runs code on your behalf, buys its most important input from a supplier you never diligenced, and can lose access to that input through an acquisition, an export order, or a court. Anysphere's customers did not choose SpaceX. Anthropic's customers did not choose the Commerce Department.

Fix the instrument first, then shorten it. And treat the answer a vendor declines to give as the most informative response on the form, because it is the only one nobody wrote for you.

Satori Canton

Founder & Principal

Satori Canton is the founder and principal of ROAI, an advisory practice focused on measuring and improving the return on enterprise AI investment.

Get a second opinion on your AI numbers.