Enterprise risk teams know how to assess a vendor. The questionnaires are long, the SOC 2 review is routine, and the approval machinery runs efficiently. When AI arrived, organizations extended those established processes to AI providers, which was a reasonable starting point and an insufficient finish line.
Vendor reviews answer one question: can this supplier be trusted? That is worth asking. The question that determines whether an AI deployment creates organizational exposure is different: can this specific system, deployed in this specific context, be trusted for the decisions it will actually influence? A vendor can satisfy every control requirement in the assessment and the deployment can still produce material harm once outputs start reaching consequential decisions. The unit of risk is the deployed system, not the company that built it.
Supplier Assurance and Deployment Assurance Are Not the Same Thing
A vendor review asks whether the provider protects data, controls access, maintains resilience, and meets contractual obligations. Those questions remain relevant when the product contains AI, and they do not establish that a specific implementation will behave appropriately once deployed inside the organization.
The same AI service can summarize internal documents, prioritize patients, recommend loan decisions, or screen applicants, each creating a materially different risk profile even though the vendor’s controls stay identical. Risk also shifts with configuration: a model limited to drafting low-stakes internal content presents a different exposure from the same model connected to sensitive records, granted access to enterprise systems, or permitted to communicate directly with customers. Autonomy, scale, data quality, and the degree to which downstream users will rely on outputs without independent verification all shape what the deployment actually risks. A system that begins as an internal drafting tool and migrates into clinical or customer-facing workflows without formal reassessment has not changed vendors. It has become a completely different deployment the original approval never evaluated.
What organizations need to assess is the complete deployed system: what decision it supports, who it affects, what the consequences of error are, how the real operating environment differs from the test environment, how human oversight actually functions under production pressure, and what happens as models, data, and business conditions change after go-live. This gap matters especially now, as AI systems approved for one purpose routinely expand into consequential processes faster than governance structures are updated to reflect that expansion. A vendor review addresses none of these questions.
Three Cases, Three Things Vendor Reviews Cannot See
The most instructive recent AI failures share a common characteristic: none involved unauthorized access or vendor security failure. All involved AI outputs reaching consequential decisions without adequate evaluation of whether those outputs were reliable, appropriate, or properly owned.
The Optum algorithm, documented in a 2019 Science study by Ziad Obermeyer and colleagues, identified patients for care management programs by predicting future healthcare spending as a proxy for medical need. Because Black patients historically incurred lower costs due to systemic barriers to care access, the model consistently rated equally sick Black patients as lower risk than white patients, reducing their identification for additional care by more than half. The researchers found that realigning the algorithm to measure illness directly would have nearly tripled the proportion of Black patients flagged for additional support, from 17.7 percent to 46.5 percent. The vendor had appropriate certifications, the data was handled properly, and the algorithm was performing exactly as designed. The failure was in the objective the model was given, and no vendor security questionnaire is built to ask whether a proxy variable is appropriate for the decision it will drive or whether the model performs equitably across affected populations.
Air Canada’s chatbot told a passenger he could apply for bereavement fares retroactively after completing his trip, which was incorrect. When he submitted the claim and the airline declined, the British Columbia Civil Resolution Tribunal found Air Canada liable for negligent misrepresentation in Moffatt v. Air Canada, 2024 BCCRT 149, rejecting the airline’s argument that the chatbot was a separate legal entity responsible for its own statements. Tribunal member Christopher Rivers called that argument remarkable, noting that a company is responsible for all information on its website regardless of whether it comes from a static page or an automated system. The governance gap had nothing to do with the chatbot vendor’s security controls. It was about who inside Air Canada owned the accuracy of statements made through an official customer channel, whether policy changes would be reflected in what the system told customers, and what monitoring would catch the discrepancy before it affected more people. Those are deployment governance questions with no home in a supplier questionnaire.
The McDonald’s automated drive-through pilot with IBM ran in more than 100 US restaurants from 2021 until it was shut down in July 2024, after the system struggled with varied accents, background noise, and mid-order changes in live environments. A supplier assessment could confirm IBM’s security infrastructure. It could not establish whether the system would perform reliably across the full range of real restaurant conditions or whether employee intervention would hold at peak service volume. McDonald’s ended the IBM partnership while affirming that voice-ordering technology remains part of its long-term plans, a signal that the specific operational context, not the technology concept, drove the decision. The lesson for any AI governance program is that approval for a promising capability must remain conditional on evidence gathered from the environment where it will actually operate, not from a controlled evaluation designed to make the system look good.
Each case reflects a different failure mode: a flawed model objective, an unowned accountability gap for what the system says on the organization’s behalf, and a deployment context that exceeded the system’s validated range. Vendor reviews surface none of them.
Human Oversight Needs an Operational Definition
Organizations frequently cite human review as a safeguard without defining what that review entails. Placing an employee nominally in the process is not a functioning control. The reviewer must have sufficient expertise, information, authority, and time to identify a problem and act on it before harm occurs.
In practice, oversight fails in predictable ways. Employees defer to systems that appear authoritative. High agreement rates produce automation bias, where nominal review becomes routine acceptance. Productivity expectations make independent evaluation impractical. Responsibility stays formally assigned to a person while the system drives the outcome.
An assessment should define oversight in operational terms: who reviews the output, what information they receive to evaluate it independently, whether they can override without penalty, how overrides are tracked, and how the organization detects when review has become rubber-stamping. Metrics like override rate, escalation frequency, and time spent on review provide early signals that nominal oversight has drifted from functional oversight. The assessment should also define circumstances where the system should not produce a recommendation at all. In the programs I have worked on, the gap between oversight that exists on paper and oversight that functions under production conditions is reliably where governance exposure accumulates.
What a Deployment-Centered Assessment Must Establish
A deployment-centered assessment starts with supplier review and extends into four areas vendor reviews leave unexamined. It asks what the system is being used to decide and what the consequences of error are, because a label like “customer support tool” does not contain enough information to evaluate actual risk. It examines the real operating context: the actual data, the real users, the environmental variability, and the incentives shaping how employees will use outputs under operational pressure. Controlled test performance does not guarantee production performance when the deployment environment differs from the one assumed during development.
It records a validity envelope for the approval: the model version, use case, affected population, data sources, and operating assumptions under which the deployment was found acceptable. It also defines what changes require a new assessment. A system approved to summarize internal documents is not automatically approved to respond directly to customers. A tool validated for one population requires performance evidence before expanding to another. An assistant authorized to make recommendations needs reassessment before receiving authority to execute transactions. Organizations should specify in advance what evidence justifies expansion: error analysis, demographic performance data, incident trends, and confirmation that human review controls are functioning as designed. This structure prevents a bounded pilot from quietly becoming a production deployment the organization never formally evaluated.
What Mature Practice Looks Like
A mature AI governance program organizes its inventory by deployed capabilities and use cases alongside vendor relationships. It links each deployment to an accountable owner, approved boundaries, monitoring thresholds, and the conditions that would trigger reassessment. Monitoring is designed during assessment rather than retrofitted after deployment: the organization decides in advance what evidence would demonstrate acceptable operation, which changes would invalidate approval, and what authority exists to restrict or suspend the system when performance degrades. Relevant indicators vary by use case but typically include error rates, override rates, demographic performance differences, escalation volume, and the rate at which reviewers identify outputs requiring correction.
Governance effort is proportionate. A low-impact internal drafting tool should not receive the same scrutiny as a system influencing healthcare, credit, or employment decisions. Concentrating scrutiny where the consequences, autonomy, and scale are greatest is what makes a mature program sustainable.
The NIST AI RMF, ISO/IEC 42001, and the EU AI Act all treat assessment as a continuous activity across the AI lifecycle. The frameworks are clear on the direction. The operational gap is the institutional habit of treating a vendor approval as a proxy for deployment assurance, which leaves the riskiest questions permanently unanswered.
A deployment-centered assessment covers five connected areas:
- The supplier and technology foundation: Security architecture, data handling, resilience, access controls, and contractual obligations. This is the layer most vendor reviews already cover well.
- The intended purpose and decision impact: What decision the system supports, who is affected, and what the organizational consequences of an error are at the use-case level.
- The deployment context: The actual data, real users, and operating conditions the system will encounter, which may differ substantially from the test environment.
- Human oversight and operating controls: Who reviews outputs, with what information, what authority to override, and how the organization detects when review has become routine acceptance rather than genuine evaluation.
- Lifecycle monitoring and reassessment: Performance thresholds tied to the risks identified during assessment, and explicit triggers for reassessment when the model, data, population, or use case materially changes.
Vendor security, privacy, and resilience reviews remain essential. A weak supplier can undermine every deployment built on top of it. The point is that a strong supplier cannot guarantee a responsible deployment, because the risks that matter most in AI governance originate in how the organization has configured the system, what it has authorized the system to influence, and whether those decisions remain sound as conditions change.
Organizations that build governance structures reflecting this distinction will be able to answer the question that actually matters: whether the deployed system remains appropriate for its approved purpose under current operating conditions. Those that do not will keep producing thorough vendor reviews while leaving their most consequential AI exposure unexamined.
About the Author
Kimly Hong is a Principal Cybersecurity and GRC Consultant with more than ten years of experience building enterprise security programs across regulated financial services, hospitality, and technology environments. Her work spans governance, risk, and compliance program design, third-party risk management, access governance, and incident response readiness. She has built these programs from the ground up across complex, multi-region environments and currently consults across financial services, SaaS, and retail organizations. Connect on LinkedIn to continue the conversation.