AI Vendor Risk Assessment: The Questions That Matter
Six questions do most of the work in an AI vendor assessment: is our data used to train your models, are prompts and outputs retained and for how long, which foundation-model providers sit behind your product as sub-processors, where does inference run, what can the system do without human approval, and can you evidence an AI governance framework. The standard instruments have caught up in form, and it is worth saying plainly because the opposite is often claimed: the Cloud Security Alliance published AI Controls Matrix v1.1 in July 2026 with 247 control objectives across 18 domains and a 320-question AI-CAIQ, and the Shared Assessments SIG carries AI content too. The problem is not absence, it is that 320 questions produce fatigue rather than discrimination, and vendors clear those reviews routinely without anyone pressing on the six that decide the risk.
This article gives the six, the follow-up that turns each into evidence, what to do with a bad answer, and where the whole exercise is not worth running.
Key takeaways
- Six questions discriminate: training use, retention, model sub-processors, inference location, autonomy, and evidenced governance.
- Standard instruments do cover AI now (CSA AICM v1.1, 247 controls, 320 AI-CAIQ questions), but length is not signal.
- Every question needs a documentary follow-up, or you have collected an assertion rather than evidence.
- Autonomy is the most under-asked and most consequential question, and it maps to a named risk category in the OWASP LLM Top 10.
- Tier by consequence, not by spend. A cheap tool touching customer data outranks an expensive one that does not.
The six questions, and the evidence behind each
1. Is our data used to train or fine-tune your models?
The follow-up is the one that matters: show me the contractual term. A yes-or-no in a spreadsheet is an assertion; a clause in the agreement or a linked, versioned data-usage page is evidence. Ask separately about training on your data and about human review of your data for quality or abuse monitoring, because those are different practices and a vendor can truthfully say no to the first while doing the second.
2. Are prompts and outputs retained, and for how long?
Retention is where regulated data quietly accumulates. Ask for the retention period, whether it differs for abuse-monitoring logs, whether you can turn retention off, and what deletion actually means operationally. Ask for it in writing, because retention defaults change between plan tiers more often than any other setting in this category.
3. Which foundation-model providers are behind your product?
Any vendor that cannot answer this immediately has not thought about its own supply chain. You need the provider names, the processing purpose, the data categories and the regions, on a stable sub-processor page with change notification. This maps directly to LLM03 Supply Chain in the OWASP Top 10 for LLM Applications (2025 edition), and it is the question that reveals whether you are onboarding one vendor or four.
4. Where does inference run?
Region matters for data residency and it matters for who can compel disclosure. "In the cloud" is not an answer. Ask which regions, whether requests can fail over to another region under load, and whether that failover is contractually bounded.
5. What can the system do without human approval?
This is the least-asked and most consequential of the six. An assistant that drafts is a different risk from an agent that sends, refunds, provisions or deletes. Ask what actions the system can take autonomously, what the blast radius of a wrong action is, and what credentials it holds. The OWASP list names this directly as LLM06 Excessive Agency, and it is the category where the gap between what a vendor's marketing implies and what the product can actually do is widest.
6. Can you evidence an AI governance framework?
Not "do you have one." Evidence means an ISO/IEC 42001 certificate, a written NIST AI RMF alignment statement, or a completed AI-CAIQ. A framework named without an artifact behind it should be recorded as no.
What the standard instruments give you, and what they cost
CSA's AI Controls Matrix reached v1.1 in July 2026, expanding from 243 to 247 control objectives across 18 domains, with the AI-CAIQ providing 320 questions mapped to those controls. It is a genuinely good artifact, it is free, and if a vendor hands you a completed AI-CAIQ you have a substantial amount of structured information.
The cost is on both sides. Sending 320 questions to a five-person vendor guarantees either a refusal or a rushed set of answers nobody will stand behind, and reading 320 answers carefully is a day of work per vendor that most programs do not have. That is why the six matter: they are the subset that changes your decision, and everything else is context you can gather later for the vendors that survive.
A workable pattern: send the six to every AI vendor, send the full instrument only to vendors in your top tier, and accept a completed AI-CAIQ in place of your own questionnaire whenever a vendor offers one, because a document they maintain is more current than one they filled in for you nine months ago.
Tier by consequence, not by contract value
The instinct is to tier vendors by spend. For AI vendors that is the wrong axis, and the reason is that the cheapest tools frequently have the deepest access.
Tier by two things instead. First, data sensitivity: does this tool see customer data, regulated data, or source code? Second, autonomy: can it act, or only suggest? A free browser extension that reads every page an employee visits and a six-figure platform that summarizes public filings are not the same risk, and a program that tiers on invoice size will review the wrong one.
Our general approach to tiering, escalation and re-review cadence is in Building a Vendor Risk Management Program; the AI-specific change is that the autonomy axis is new and most existing programs do not have a field for it.
Where the regulatory hooks actually bite
Two points worth being precise about, because both are frequently overstated.
If you deploy a high-risk AI system in the EU, the EU AI Act places duties on you as a deployer, not only on the provider. Those duties are on a moving calendar: Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force on 27 July 2026, deferred the full high-risk obligations to 2 December 2027 for standalone Annex III systems and to 2 August 2028 for AI embedded in Annex I regulated products, while the Article 5 prohibitions, general-purpose AI rules and Article 50 transparency duties kept their original schedule. Scoping matters more than the date: most SaaS tools are not high-risk systems, and we work through the classification in Understanding the EU AI Act: What US Companies Need to Know.
Second, US state privacy law increasingly reaches automated decision-making and profiling through opt-out rights and assessment duties, which means a vendor's autonomy answer can create an obligation for you rather than for them. Check which state regimes you are already subject to before assuming an AI tool is out of scope.
What neither of these creates is a general legal requirement to run an AI vendor questionnaire. You run it because a bad AI vendor is a live operational risk, and because your own customers are asking you the same questions.
What to do with a bad answer
Most programs collect answers and file them. The value is in what happens next, and there are only four honest options.
- Accept the risk, in writing, with a named owner and a review date.
- Mitigate it, by restricting the data the tool sees, disabling retention, or removing the tool's ability to act rather than suggest.
- Contract around it, with a data-usage term, a sub-processor notification clause, or a deletion commitment.
- Decline the vendor.
The failure mode is a fifth option nobody writes down: note the gap, buy the tool anyway, and never revisit it. If your assessment cannot produce a decision, it is documentation theater, and we would rather you ran the six questions and acted on them than ran 320 and filed them.
When this is not worth doing
If the tool never touches customer data, never touches regulated data, cannot take an action, and would cost you nothing to switch away from, the six questions are enough on their own and a formal assessment is overhead. Some vendors will encourage you to assess everything; that recommendation is easier to make when it is not your calendar.
If you have fewer than about twenty vendors in total, a tracked spreadsheet and the six questions beat a purchased platform. Tooling helps when the volume is real; below that, it mostly moves the same neglect into a nicer interface.
And if nobody in your organization owns vendor risk, buying an AI questionnaire changes nothing. Ownership first.
Where Top Floor fits
We build the AI extension into an existing vendor risk program rather than selling a separate AI program: the six questions in the intake, the autonomy field in the tiering model, and the escalation path for the answers that come back badly. Where there is no program to extend, our vCISO engagements stand one up at a size the company can actually run.
For companies where the volume is continuous, compliance as a service covers the re-review cadence, which is the part that decays: model providers change, vendors add autonomy to products that used to only suggest, and last year's answers stop being true without anybody sending a notice.
How to decide this week
- List every AI tool in use, including the ones bought on a card and never told to IT.
- Add one field to your vendor record: can it act without human approval, yes or no.
- Send the six questions to your top-tier AI vendors and ask for the contractual term behind each answer, not the assertion.
- Check whether your existing data processing agreements require sub-processor notice, and whether your vendors have been giving it.
- Pick one vendor whose answers come back badly and take an actual decision on it, so the process proves it can produce one.
Frequently asked questions
What questions should I ask an AI vendor?
Six discriminate: is our data used to train or fine-tune your models, are prompts and outputs retained and for how long, which foundation-model providers sit behind your product as sub-processors, where does inference run, what can the system do without human approval, and can you evidence an AI governance framework with a certificate or a written alignment statement. Ask for the contractual term or the linked policy behind each answer rather than accepting a yes in a spreadsheet cell, because an assertion you cannot point at later is not evidence.
Do standard security questionnaires cover AI risk?
Yes, and claims to the contrary are out of date. The Cloud Security Alliance published AI Controls Matrix v1.1 in July 2026 with 247 control objectives across 18 domains and a companion 320-question AI-CAIQ, and the Shared Assessments SIG carries AI content as well. The practical issue is length rather than coverage: 320 questions produce fatigue on both sides, so a workable pattern is to send a short discriminating set to every AI vendor, reserve the full instrument for your top tier, and accept a vendor's own completed AI-CAIQ when they maintain one.
How should I tier AI vendors?
By consequence rather than by contract value. The two axes that matter are data sensitivity (does the tool see customer data, regulated data, or source code) and autonomy (can it take actions, or only suggest them). Tiering on spend reviews the wrong vendors, because a free browser extension with access to every page an employee visits can carry more risk than an expensive platform that only reads public data. Autonomy is the newer axis and most existing vendor risk programs have no field for it yet.
What if an AI vendor gives a bad answer?
There are four honest responses: accept the risk in writing with a named owner and a review date, mitigate it by restricting data access or disabling retention or removing the tool's ability to act, contract around it with a data-usage or deletion or notification term, or decline the vendor. The failure mode is a fifth option nobody documents, which is to note the gap, buy the tool anyway and never revisit it. An assessment that cannot produce a decision is documentation rather than risk management.
Related Services
Need help with your compliance program?
Our team of senior practitioners can help you navigate complex compliance requirements and build a security program that holds up under scrutiny.
Schedule a Free ConsultationGet insights like this in your inbox
Practical compliance and security guidance for teams preparing for their next audit. No spam, unsubscribe anytime.
Ask to be added to our mailing list for practical compliance and security guidance. We add you by hand, we confirm before sending anything, and we never share your address.