Two AI vendors go through the same enterprise security review. Both hand over clean SOC 2 Type II reports from established firms. The buyer’s reviewer opens the second report and finds something the first one lacks: a system description that names model lineage, inference logging, and drift monitoring, with controls tested against each. Same criteria, same opinion, two completely different audits.

This is not an auditing failure in the usual sense. It is what happens when a standard written in 2017 meets a technology that rewrote enterprise procurement years later. The AICPA has published no AI-specific Trust Services Criteria: no criteria, no points of focus, no addendum. Auditors who think an AI feature deserves scrutiny have to improvise what to ask for. Some ask for a lot. Some never start. Nothing in the written standard settles the question, and that is why the badge no longer tells a buyer what it used to.

The criteria are older than your model

Every SOC 2 report tests against the 2017 Trust Services Criteria, with a points-of-focus refresh in 2022. The refresh adjusted wording and applicability; it did not confront machine learning, let alone models that generate text and code on demand. The Cloud Security Alliance’s August 30, 2026 research note on what it calls the SOC 2 AI gap makes the timeline plain: the operative criteria were written “well before generative AI reshaped enterprise software procurement.”

Be clear about what still works. The security, availability, and confidentiality criteria apply to AI systems the way they apply to any system: perimeter, access management, change management, and encryption around the model infrastructure all get tested. A SOC 2 covering an AI product is not meaningless. What’s missing is everything specific to the model itself: training-data governance, drift, bias, non-determinism. Examination firms including Baker Tilly and Schellman have said openly that the criteria were never designed for those risks.

Nothing in the pipeline closes the gap today, either. Per the CSA note, there is no exposure draft and no task force proposing AI criteria. The closest artifact the AICPA has produced is guidance from early 2026 on responsible AI use in forensic and valuation services engagements, guidance that explicitly disclaims status as authoritative or a standard, and does not address SOC 2 engagements at all. Inside the profession the question is at least being asked: the Arizona Society of CPAs ran an expert Q&A on September 10 debating where AI belongs in the SOC 2 internal controls framework, including whether it merits a sixth category alongside security, availability, processing integrity, confidentiality, and privacy. Debated, not decided.

What auditors are actually asking for

Without a written baseline, examination firms improvise, and their improvisations are converging on four evidence categories. The CSA note catalogs them, and a September 4 synthesis of current fieldwork published by SOC2Auditors.org describes the same four leading the request lists at firms including Baker Tilly, Schellman, and Linford & Co:

  • Model lineage. The training dataset snapshot, code commit, hyperparameters, and the approval that promoted the model to production: the record that lets anyone reconstruct what actually produced a given output.
  • Per-inference logging. Model version, redacted prompt or prompt hash, tool calls, and outcome captured for each inference: the trail incident responders need when an AI feature is part of an incident.
  • LLM subprocessor documentation. If your product calls an external model provider, that provider becomes your customer’s subprocessor too. Auditors want its own SOC 2 or ISO 27001 status and its data-retention configuration.
  • Drift monitoring tied to change management. Dashboards watching model behavior over time, paired with change tickets that show the model running in production is the model that was tested.

None of these appears in any AICPA criterion. That is the whole problem in one sentence: an auditor can demand all four, another can close the engagement without requesting any, and both reports come back clean.

Two clean reports, two different audits

For buyers, the consequence is direct. Whether model accuracy, training-data integrity, or drift detection got tested depends on which auditor was engaged and which practitioner overlay that firm happens to use. That is CSA’s phrasing, and it matches what shows up in vendor files we review. “SOC 2 Type II” is one signal, not a verdict.

The opinion page is the least informative part of the report. It attests that the presentation is fair against whatever was scoped; it does not attest that the scope was adequate. The system description and management assertion are where scope actually lives. If model lineage and drift monitoring never appear there, they were not tested, whatever the overall cleanliness of the report suggests.

The quality squeeze

While the criteria sit still, the AICPA is tightening scrutiny of how SOC 2 reports get produced. In February, the Journal of Accountancy reported the Assurance Services Executive Committee’s SOC 2 Working Group warning that the flood of “fast and easy” attestation platforms is putting the credential’s credibility at risk: “[SOC] professionals are seeing indications that ‘fast and easy’ may come at the expense of quality and objectivity,” said Sean Linton, the working group’s chair and an audit partner at EisnerAmper.

In May, the AICPA Peer Review Board went further, issuing a reviewer alert that directs added scrutiny to firms with SOC 2 practices beginning June 1, 2026, flagging identical reports, risk assessments, sample sizes, and testing procedures across clients as potentially nonconforming engagements.

Put the two together and the dynamic gets awkward. A firm that standardizes one AI overlay and applies it uniformly across clients looks a lot like the pattern peer reviewers are now hunting for. A firm that improvises per engagement produces reports that diverge, which is why buyers can’t compare them in the first place. Neither path yields comparability, and neither is exactly wrong; the written baseline that would reconcile them doesn’t exist. If you are commissioning a report, treat auditor selection as a scope conversation first and a price conversation second: the differences that matter now are the evidence categories a firm will actually commit to testing. We keep separate notes on why SOC 2 alone is not a security strategy; the scoping discipline is the same for AI-heavy systems.

If you sell an AI product

Two moves, and both get more expensive after fieldwork starts.

First, the system description. It defines the audit boundary, and controls not described there are generally not tested or reported as tested. CSA’s advice to companies commissioning reports is to make the engagement partner name the AI evidence categories explicitly in the system description before examination work begins. Saying it in scoping costs one conversation. Discovering the gap after fieldwork means a thinner report than your buyers expect, or a mid-examination scramble to widen scope.

Second, the evidence itself. The four categories double as a build list: a model registry that captures lineage at promotion time; inference logging with prompt redaction designed in rather than bolted on; a subprocessor file for every external model API you call, with report status and retention configuration; drift dashboards whose alerts open change-management tickets, so monitoring and change control tell one story. Teams that wait for the auditor’s request list end up assembling all of this under deadline, and thin evidence is what that produces.

If you buy AI

Add the four categories to your vendor questionnaire as named items and score the answers instead of filing them. Expect wide variance for now, because the variance is information. A vendor whose system description names drift monitoring has materially more mature AI governance than a vendor whose report never mentions AI at all, regardless of what both opinion letters say.

Then actually read the report you’re given: system description first, management assertion second, opinion letter last. If none of the four categories appear anywhere, you have learned something about that vendor’s AI governance, and it is not “compliant.” We’ve written before about a security question that killed a term sheet. Diligence questions that produce silence are cheaper asked than discovered post-signature, and an AI evidence section is the same idea pointed at 2026 procurement.

Interim options, with fine print

CSA publishes the AI Controls Matrix and runs STAR for AI, and its note positions them as the bridge until official criteria exist. The mapping is serious work and can be useful when you structure a scope conversation with an auditor. But read CSA’s own caveat first: “independent validation of that mapping inside live SOC 2 engagements remains limited.” Treat AICM and STAR for AI as one option among several for organizing the discussion, not as a credential, and not as a substitute for criteria that don’t exist.

The same evidence serves three frameworks

The four pipelines are not just audit armor. NIST AI RMF’s Measure and Manage functions ask for exactly this: measurement of model behavior and evidence that controls operate. ISO/IEC 42001’s performance-evaluation clauses want monitoring records too. One instrumentation build, citable against three frameworks. And it starts with inventory: none of this is producible for AI systems nobody has inventoried, which is why we treat AI inventory and shadow-AI discovery as the first milestone rather than a cleanup task.

Two boundary notes. Teams selling into the EU should know the AI Act’s post-market monitoring obligations push toward the same drift telemetry. We covered Articles 72 and 73 separately, so this piece stays on the US attestation question. And per-inference logging does double duty in incident response: for teams with CIRCIA obligations, the 72-hour forensics clock is easier to satisfy when the trail already exists.

What to do next

If you own a SOC 2 and ship an AI feature, the work is concrete:

  • Sellers: at your next scoping conversation, ask the engagement partner to name AI evidence categories (lineage, inference logging, subprocessor documentation, drift monitoring) in the system description. Build the four pipelines before fieldwork, not during.
  • Buyers: rewrite the AI section of the vendor questionnaire around the four categories, and start reading system descriptions instead of opinion pages.
  • Either side: there is no written baseline to wait for as of today. Organize around the evidence auditors actually request, and document your scope decisions so your next report is comparable to this one.

None of the four pipelines is exotic. Most of the work is deciding what “good” looks like for your model, then wiring telemetry you may already half-have. That gap analysis and build-out is the work NTD’s AI governance practice does with mid-market fintech, crypto, and SaaS teams, including the scope conversation with your auditor. If you would rather walk into fieldwork with the evidence categories already named and the pipelines already running, start there.