Est.

Explainability Requirements in AI Assurance

Auditors now require AI systems to prove they can explain their decisions.

Features Editor · · 10 min read
Cover illustration for “Explainability Requirements in AI Assurance”
AIUC-1 and AI Assurance · September 25, 2026 · 10 min read · 2,205 words

Explainability used to be in the "nice to have" pile for AI systems, the kind of thing a product team gestures at in a slide about responsible design. Not anymore. Across financial reporting, service organization audits, and a new wave of AI-specific certification schemes, auditors now test explainability as a line item, with logs, disclosures, and documentation that either exist or don't. The gap between those two states is where most of the exposure sits right now, and it runs wider than most compliance teams realize.

Start with the mismatch driving all of this. Survey work in the space has repeatedly found that a large majority of consumers, on the order of 85%, want an explanation when an AI system makes a decision that affects them. Meanwhile, a good chunk of production models, estimated around 60% in industry research, still run as unexplainable black boxes. That gap used to be a reputational risk. It has become a legal one, because regulators stopped treating explainability as aspirational language and started writing it into binding text. EU data protection rules already require explanations for automated decisions that affect individuals. The EU AI Act goes further, requiring transparency documentation for high-risk systems, with the standalone high-risk obligations addressed through the Digital Omnibus agreement. Once explainability sits in statute, it is no longer philosophy but a checklist item somebody has to satisfy, on a deadline, with evidence.

Financial statement audits absorbing AI explainability into core expectations

Financial reporting is where this shift turns concrete fastest, because financial statements get audited against standards that already assume someone can explain how a number got there. When an AI model produces an estimate that feeds a material line item, its audit trail has to survive the same scrutiny as a spreadsheet formula or an actuarial model. Nobody gets to wave a machine learning model through on the strength of "it's usually right," and that's the assumption a lot of engineering teams still carry into their first audit.

Four roles inside a finance organization actually have to act on this, and each one owns a different piece of the problem. A CFO needs to know what disclosure obligations attach to AI use in the reporting process. A controller has to work out which AI applications touch something material enough to require disclosure at all, since not every automated tool clears that bar. Internal auditors need to fold AI into their testing of internal control over financial reporting (ICFR), treating a model the way they'd treat any other system that touches a number on the balance sheet. Audit committees need enough visibility into AI governance to sign off that the oversight is adequate.

The PCAOB's 2025 inspection priorities gave this weight by flagging generative AI use at public companies and broker-dealers as an area of heightened inspection emphasis. That's a regulator telling audit firms to look harder at how AI shows up in the numbers they sign off on.

PCAOB AS 1201 isn't a new standard, but it's being pointed at a genuinely new kind of tool now. Both carry the same underlying logic: when an auditor relies on the work of a specialist, such as a valuation expert or an AI system, the auditor still owns the conclusion. Ownership requires understanding, full stop. An AI-assisted risk assessment that spits out a number nobody on the audit team can walk back through and explain fails that standard, no matter how accurate the model turns out to be. Accuracy without explainability doesn't cut it here. The auditor's signature is a claim of understanding.

What SOC 2 covers when AI is involved

SOC 2 remains one of the most requested attestations in security, certifying controls at a service organization across a defined set of trust service categories as they apply to systems that store or process customer data. That scope is real and useful. Buyers often assume AI gets covered the moment the word enters the conversation, and that assumption is where the trouble starts.

When AI shows up inside a SOC 2 audit, the auditor checks things like whether access to model training data is properly restricted (a security control) or whether backups exist for the systems running the model (an availability control). Those checks matter. They tell you nothing about how the model actually behaves.

Model bias, machine learning explainability, hallucination behavior, resistance to prompt injection, and AI-specific governance structures all sit outside SOC 2's defined scope, and no amount of wishful reading of the report changes that. A SOC 2 report can confirm a vendor's infrastructure is reasonably locked down. It can't say whether that vendor's AI agent resists a prompt injection attack, leaks one customer's data into another customer's session, or invents a policy answer under pressure rather than admitting it doesn't know. Enterprise buyers have started asking exactly those questions, questions SOC 2 was never built to answer.

Treating a SOC 2 report as proof of safe AI behavior is already a mistake. The framework does what it says, no more, and the boundary sits closer than most procurement teams think it does.

How AIUC-1 structures explainability as an auditable requirement

AIUC-1 was built to sit past that exact boundary. Created by the Artificial Intelligence Underwriting Company (AIUC), the standard developed with input from Orrick, MITRE, the Cloud Security Alliance, Stanford, and a consortium of enterprise security leaders. More than 120 consortium members contributed to the Q2 update through technical sessions and peer review. That's a lot of enterprise weight sitting behind a standard barely out of its first year.

The structure is granular on purpose: 51 requirements broken into 130 controls, split between 65 mandatory and 65 optional, organized across six pillars: Data & Privacy, Security, Safety, Reliability, Accountability, and Society. How many controls actually apply depends on what the agent does. A narrow, single-purpose agent typically needs to satisfy around 40 controls. A complex, multi-modal agent operating across more surfaces needs closer to 65.

Explainability lives inside the Accountability pillar, and three requirements carry the weight. E015, mandatory, requires logs of AI system processes, actions, and agent outputs sufficient to support incident investigation, auditing, and after-the-fact explanation of system behavior. E016, also mandatory, requires clear disclosure mechanisms so users know they're interacting with an AI system rather than a human. E017, optional, asks organizations to maintain a system transparency policy along with a repository of model cards, datasheets, and interpretability reports for major systems.

The mandatory pair forms a floor nobody gets to skip. Log the behavior, disclose the interaction, or the deployment doesn't count as accountable, no matter what else the agent gets right.

How AIUC-1 compares to ISO 42001 and SOC 2

Lining these three up, the cleanest way to keep them straight is to ask what question each one is actually built to answer. They aren't competing for the same job, and treating them as substitutes for one another is where most confusion starts.

SOC 2 asks whether a service organization runs sound security controls. It has nothing to say about how an AI system behaves once those controls sit in place. ISO 42001 asks a different question: does the organization run a responsible AI management system, meaning governance structures, defined processes, a mechanism for continual improvement? ISO 42001 certifies the organization as a whole, not any single agent running inside it. AIUC-1 asks something narrower and more operational: does this specific AI agent perform safely, securely, and accountably under real and adversarial conditions? That combination of governance assessment and technical evaluation of the agent itself gives it a different unit of analysis than either of the other two frameworks uses.

The relationship between AIUC-1 and ISO 42001 runs close but not interchangeable, and AIUC's own published crosswalk says so directly. AIUC-1 incorporates many of ISO 42001's controls, but AIUC's own published crosswalk identifies areas where full coverage gaps and partial overlaps remain. Practically, certifying against AIUC-1 doesn't satisfy ISO 42001, and holding ISO 42001 doesn't satisfy AIUC-1 either. That looks like redundant paperwork until you look at what each one actually verifies.

ISO 42001 builds the management system foundation, the policies and governance scaffolding an organization runs on. AIUC-1 tests whether those policies produce effective controls in a live, running agent under adversarial pressure. One is a blueprint. The other is a stress test. Organizations that hold ISO 42001 may still pursue AIUC-1 for exactly that reason, because paperwork isn't behavior, and only one of these frameworks tests behavior directly.

The audit cadence tells the same story from a different angle. SOC 2 and ISO 42001 run on annual or otherwise fixed audit cycles, which fits how those frameworks think about organizational controls. AIUC-1 audits quarterly. Agent architectures shift fast enough that a once-a-year check would miss real changes in the threat surface between audits. Recent AIUC-1 update cycles have added requirements covering MCP and A2A protocol security, non-human identity management, and agent access controls, categories that barely existed as named risks two years earlier.

What the early adopters reveal about who this certification is for

AIUC-1 launched in July 2025. Schellman became the first accredited auditor under the standard, and ElevenLabs became the first voice AI company to earn certification, an early signal of where initial demand concentrated.

The more telling milestone landed in March 2026, when UiPath became the first enterprise automation platform to certify against AIUC-1, putting its agentic systems through more than 2,000 technical evaluations in the process. UiPath's involvement isn't incidental: the company was also a founding technical contributor to the standard itself. It helped write the bar it then had to clear.

Lining up the early cohort reveals a pattern fast. Voice AI, enterprise automation, workflow platforms: these are categories where the agent touches high-volume, high-stakes transactions, and where a failure carries real operational and reputational cost rather than a mild inconvenience. That's not coincidence: it's a reasonable early-adopter profile for a certification built around agent behavior under adversarial conditions, because the companies with the most to lose from an unexplainable failure are the ones adopting it first. It's a reasonable early-adopter profile for a certification built around agent behavior under adversarial conditions, because the companies with the most to lose from an unexplainable failure are the ones showing up first to prove they can explain one.

AI-related questions currently make up a small share, roughly 3%, of items on a typical enterprise security questionnaire. Small, yes, but the content of those few questions is already narrow and specific: hallucination risk, prompt injection defenses, unauthorized agent actions. That's precisely the terrain AIUC-1 was built to cover, and the standard emerged roughly in step with the questions buyers had already started asking on their own.

What organizations must document and demonstrate to satisfy explainability expectations in practice

What does an organization actually need on file when an auditor comes asking? Across SOX, HIPAA, FFIEC, PCI DSS v4.0, and the EU AI Act, the documentation requirements converge on a defined minimum schema for AI activity logs. Finance organizations tend to hit the same three gaps first, and in the same order: no clear attribution when a service account, rather than a named human, took the action; confidence scores sitting in a log where actual reasoning should be; and retention windows that fall short of what SOX requires.

Under AIUC-1, the mandatory floor is narrower but not negotiable. E015 requires logs of AI system processes, actions, and agent outputs sufficient to support incident investigation and auditing. E016 requires clear disclosure mechanisms so users know they are interacting with an AI system rather than a human. Neither one sits as aspirational language in a policy binder somewhere collecting dust. Both are controls an auditor tests directly, with evidence, on a quarterly cycle, which is a shorter leash than almost any other compliance regime in this piece.

The word "explanation" gets used loosely, doing three different jobs at once. Process-level explanation covers what the agent did, in what order, touching what data. Decision-level explanation covers what inputs produced what output, and with what confidence attached. Governance-level explanation covers who authorized the agent's scope in the first place, what policy governs its behavior, and who owns the response when something breaks. A system can satisfy one of these levels and fail the other two outright, so frameworks like AIUC-1 test them as separate controls instead of folding everything into one generic explainability checkbox.

None of this gets fixed by writing better logging code. The real problem sits one layer down, in identity. AI agents often run under service identities with no named human attached to the action, which means the audit trail can show what happened without ever being able to say who, or what, was accountable for it. That's a structural gap, not a documentation gap, and closing it takes deliberate identity architecture built for non-human actors from the ground up. Frameworks like AIUC-1 are only just beginning to formalize that work into testable, mandatory controls, which suggests the current mandatory floor, E015 and E016, is a starting point rather than the finished standard. Anyone treating those two controls as the ceiling of what explainability requires is going to find that out the hard way, probably during an incident review rather than an audit.

Diagram: The Explainability Gap: 85% Want It, 60% of Models Can't Deliver It. Visualizes: Show the contrast between two numbers from industry research: 85% of consumers want an explanation when an AI decision affects them, versus roughly 60% of…

Sources

  1. What Is AIUC-1? The First Security Standard Built for AI Agents | Workstreet
  2. vettedaiagents.com
  3. britive.com

More in AIUC-1 and AI Assurance