Est.

AIUC-1 Standard for AI System Audits

Overhauls how enterprises audit AI agents that make decisions autonomously.

Staff Writer · · 13 min read
Cover illustration for “AIUC-1 Standard for AI System Audits”
AIUC-1 and AI Assurance · September 12, 2026 · 13 min read · 2,980 words

Agentic AI systems don't just answer questions. Given a goal, they break it into subtasks, pick tools, take action in live systems, and check their own work, sometimes across hundreds of decision cycles before a human ever sees the output. AIUC-1 is the first standard built specifically to audit that kind of behavior, and the bet this piece makes is a direct one: the standards enterprises have leaned on for a decade, SOC 2 and ISO 27001 chief among them, cannot do this job, and pretending they can is the mistake currently costing companies real money.

The gap isn't abstract. An agent that books a flight, refunds a customer, or rewrites a database entry has changed something real in the world, and that's a lot harder to undo than a bad chatbot reply. Nobody can say for certain who "owns" a decision an agent made on its own, so accountability gets fuzzy fast. And because an agent's behavior shifts depending on what tools it can reach and how the systems around it respond, risk isn't fixed at deployment. It moves. A risk assessment done in January can be stale by March, not because anyone did anything wrong, but because the agent's environment changed underneath it while nobody was looking.

SOC 2 and ISO 27001 were built for static software and predictable data flows, long before agentic systems became something enterprises actually ran at scale. Neither one checks whether a live agent resists a jailbreak attempt or leaks sensitive data under adversarial pressure, and no amount of goodwill from either standard's authors changes that gap. The stakes are concrete, not hypothetical: EY found that 64% of companies with over $1 billion in revenue have already lost more than $1 million to AI-related failures, and Cisco's 2025 AI Readiness Index found only 29% of companies believe they're actually equipped to defend against AI-specific threats. IBM's 2025 Cost of a Data Breach Report sharpens the picture further: 97% of organizations that had an AI-related security incident lacked proper access controls for their AI systems in the first place. Meanwhile, over 95% of enterprise AI pilots reportedly never make it to full deployment, and legal and security concerns get cited most often as the reason they stall out. So the question worth sitting with is whether a new standard was needed at all is not really the point. It's why it took this long.

What AIUC-1 is and where it came from

AIUC-1 stands for Artificial Intelligence Unified Controls, version 1. It's published by the Artificial Intelligence Underwriting Company, a San Francisco startup that came out of stealth in July 2025 with a $15 million seed round led by Nat Friedman at NFDG, joined by Emergence, Terrain, and angel investors including Anthropic co-founder Ben Mann and former CISOs from Google Cloud and MongoDB.

The founding team's background matters here, and it's worth being specific about why. CEO Rune Kvist was an early product and go-to-market hire at Anthropic. Brandon Wang, a Thiel Fellow, had already built a consumer underwriting business before this. Rajiv Dattani came out of McKinsey and served as COO of METR, the AI evaluations nonprofit. Between the three of them, they've sat inside both the frontier labs building these systems and the institutions trying to price and govern the risk those systems create. That combination, lab experience plus underwriting experience, is the whole thesis of the company in miniature.

AIUC didn't write this standard alone in a room. It came out of a consortium of more than 100 Fortune 500 CISOs and security leaders, with technical input from Stanford's Trustworthy AI Research Lab, MIT, MITRE, Cisco, Google Cloud, and the Cloud Security Alliance. Orrick, the law firm, contributed too, which explains why the standard covers legal and regulatory exposure and not just technical vulnerabilities.

Buyers tend to miss one distinction, and it's the one that matters most: AIUC-1 certifies a specific AI system or product, not an entire company. A vendor could have one agent certified and three others that aren't. Assuming a badge on a vendor's homepage covers everything they sell is exactly the kind of mistake this standard was built to catch. Checking scope before signing anything isn't optional, and any procurement team that skips this step is buying a badge, not a guarantee.

The most unusual part of the model, though, is what AIUC pairs with the audit: actual insurance, underwritten through Lloyd's of London. Systems that pass a more rigorous audit get better insurance terms. The certification becomes the evidence layer, and the policy becomes the money behind it. That's a genuinely different structure from a typical compliance badge, which gives a company a stamp and nothing else.

The six risk pillars and 130 controls that make up the standard

AIUC-1 runs on 51 requirements and 130 controls, spread across six risk pillars. Of those 130 controls, 65 are mandatory and 65 are optional, and which ones apply depends on what the system actually does. A narrow, single-purpose deployment might only need to clear around 40 controls. A complex, multi-modal agent, one juggling voice, tool use, and sensitive data all at once, is looking at closer to 65 controls.

The six pillars break down like this. Data & Privacy covers PII leakage, cross-customer data isolation, IP protection, and safeguards against sensitive enterprise data being improperly accessed or exposed. Security covers prompt injection defense (the attack where instructions hidden inside a document or webpage trick an agent into dropping its original task), adversarial robustness testing, and prevention of unauthorized actions. This pillar maps to MITRE ATLAS and the OWASP Top 10 for LLM Applications, both of which AIUC-1 threads through several pillars rather than treating as one checkbox.

Safety covers harmful output prevention, pre-deployment testing, and a documented risk taxonomy, and it specifically requires third-party testing for high-risk output categories. A vendor doesn't get to self-certify its way past this one, and that's not a small detail: it's the difference between a standard with teeth and one without. Reliability targets hallucination prevention and restricts unsafe tool calls, the failure mode where an agent takes a real action based on a premise it fabricated or misread entirely. Accountability requires clear human ownership over AI decisions, a documented failure response plan for breaches and harmful outputs, vendor due diligence, and disclosure obligations. Society is the broadest pillar, aimed at guardrails against systemic risks like cyber exploitation, CBRN misuse, and threats to national security.

Together, the six pillars trace the whole lifecycle of a deployment: how the use case gets defined, how risks get flagged up front, how the system gets watched once it's live and making its own calls. What actually separates this from a policy checklist sitting in a compliance folder is that every control has to be assessed, evidenced, and technically tested. Writing it down isn't enough, and any organization treating this as a documentation exercise has already misread the assignment.

How the certification process works in practice

Certification typically runs 4 to 8 weeks, depending on how mature an organization's AI governance already is going in. The process moves through four phases.

Scoping and kickoff comes first: nailing down exactly which agent, which tools, and which data flows are under review. Organizations running multiple agents get pointed toward the highest-risk deployments first, meaning the ones with the most autonomous decision-making power, the most sensitive data access, or the ugliest consequences if something breaks.

Next comes gap assessment and evidence collection, where the organization pulls together its policies, technical documentation, and existing testing procedures, and AIUC works alongside them to close whatever holes turn up. Then technical testing: AIUC runs live evaluations against the deployed system, probing for hallucinations, unsafe tool calls, prompt injection, data leakage, and harmful outputs, using adversarial scenarios pulled from current AI security research, including active jailbreak techniques circulating right now. ElevenLabs' certification process, for one, involved more than 5,000 adversarial simulations spread across all six pillars, which gives some sense of just how far this testing actually goes. That number is worth sitting with for a second: 5,000 simulated attacks against a single deployed system is not a spot check, it's closer to a stress test run until something breaks or nothing does.

Finally, Schellman, acting as the independent assessor, checks the evidence against the standard, and the certificate, a full audit report, and a badge are then issued. Schellman's role here is that of an independent evaluator, assessing organizations against the standard rather than acting as a remediation partner. That separation is what makes the findings mean something, because an assessor with a consulting relationship to the organization has every incentive to grade generously, and one with no such relationship has none.

The certificate itself lasts 12 months, but quarterly technical re-testing is required to keep it valid. So the badge on a vendor's website reflects current status, not a snapshot from whenever they happened to pass. It's a running claim, re-checked four times a year, not a trophy on a shelf.

Why AIUC-1 updates quarterly instead of annually

Most compliance standards update annually, if that. AIUC-1 updates every quarter, and the reasoning is straightforward once you think about how fast AI threats actually move: a set of adversarial techniques current in January can be outdated by summer. Documentation certified against last year's standard risks becoming a stale artifact instead of a live picture of risk. That's the whole bet this cadence is making, and it's arguably the single most defensible design choice in the entire standard, because an annual cadence applied to agentic AI would be a standard permanently describing last year's threats.

Each quarterly update comes out of collaboration between AIUC-1's Technical Contributors, the broader Consortium, and outside peer reviewers, so it's not AIUC unilaterally deciding what matters next. The Q1 2026 update, released in January, modified 26 requirements and added more than 40 new voice-specific controls, a direct response to how fast voice-enabled agents have scaled. The Q2 2026 update, effective April 15, 2026, modified 14 requirements and added 23 new controls, with a heavy focus on Model Context Protocol (MCP) security, plus new requirements around secrets management and secure code generation for coding agents.

The MCP focus deserves a beat of attention. MCP is the protocol increasingly used to connect agents to external tools and data sources, and it's a genuinely new attack surface, one that simply didn't exist when SOC 2 or ISO 27001 were written. A framework that only updates once a year has no real way to address a risk category that didn't exist twelve months earlier. Why would it? The standard would already be a year behind the thing it's supposed to be checking.

That's the real contrast with most security audits: they look backward, confirming what already happened. AIUC-1's quarterly cadence is trying to look sideways and slightly forward, toward whatever the next adversarial technique turns out to be. Whether that pace is sustainable once the standard covers thousands of certified systems instead of a handful is a separate question, and one the standard hasn't had to answer yet.

How AIUC-1 relates to ISO 42001, NIST AI RMF, and the EU AI Act

ISO 42001 asks whether an organization has the right policies, risk assessments, and governance structure around AI. It's a management-system audit, concerned with paperwork and process, not with whether a specific deployed agent can be tricked into leaking a customer's data. AIUC-1 mirrors some of that structure but goes further, adding controls that get audited and technically tested instead of just written up somewhere. Notably, AIUC-1 and ISO 42001 are built to sit alongside each other rather than compete for the same claim. A company holding both is making two separate arguments: here's how AI gets governed on paper, and here's how the AI actually behaves under adversarial pressure. Neither one substitutes for the other, and treating them as interchangeable misses the point of having both.

NIST AI RMF is a similar story in a different register. It gives high-level guidance organized around broad governance and risk management functions, but it doesn't prescribe specific controls or testing procedures. AIUC-1 essentially translates those functions into something that can actually be audited, with defined evidence requirements attached to each one.

The EU AI Act is where the regulatory clock starts ticking. Enforcement begins August 2026, and AIUC-1 already maps to key EU AI Act obligations. So an organization going through AIUC-1's safety and accountability pillars is, more or less as a side effect, building the evidence base it'll need for EU compliance anyway, rather than starting that work cold later. AIUC-1 also plugs into IBM Research's AI Risk Atlas Nexus, giving organizations a documented path from identifying a risk to actually fixing it, which ties the standard into risk infrastructure companies are already using.

Put simply, AIUC-1 behaves like a control overlay, not a replacement. It sits on top of frameworks a security leader is probably already running, adding the agent-specific layer those frameworks were never built to cover. Anyone pitching it as a substitute for ISO 42001 or NIST AI RMF has misunderstood what it's for.

Early certifications and what they signal about adoption trajectory

Schellman became the first accredited AIUC-1 auditor in February 2026, and that matters because Schellman already has a long track record running SOC, FedRAMP, and ISO examinations. This is an assessor already fluent in the job from the outset. It's an established one extending into new territory, which lends the early certifications more weight than a brand-new auditing body could have offered.

ElevenLabs became the first voice AI company certified under AIUC-1, and on February 11, 2026, it became AIUC's first policyholder, with its agentic products backed by a $50 million Lloyd's of London policy. That's a serious amount of insurance capital to put behind a single certification. ElevenLabs' ElevenAgents platform powers more than 3 million voice agents deployed across enterprises, and its technology reportedly runs somewhere inside more than 75% of Fortune 500 companies, including Cisco, Square, Revolut, and MasterClass.

Then in March 2026, UiPath became the first enterprise automation platform to get certified, a signal that adoption is moving past AI-native startups and into mainstream enterprise software, which is the harder market to crack. UiPath built its AIUC-1 certification on top of an existing ISO/IEC 42001:2023 certification, exactly the layered approach the standard is designed around. UiPath is also a Founding Technical Contributor to AIUC-1, so it helped write the rules it's now being measured against, a detail that deserves attention rather than being glossed over.

The pattern here looks a lot like how SOC 2 became standard practice, and that comparison isn't incidental. SOC 2 never became mandatory through legislation. It became mandatory because enterprise buyers started demanding it, and vendors without it started losing deals to vendors with it. AIUC-1 looks to be tracing the same arc, at a moment when AI trust failures make headlines often enough that buyers are paying attention. A recent Workstreet analysis found that roughly 3% of enterprise security questionnaire questions now touch on AI specifically. That's a small slice today. But given how fast agentic deployments are scaling, betting that share stays flat is the wrong bet, and any vendor planning around a flat trajectory is planning around the wrong number.

What AIUC-1 certification means for organizations on both sides of the buyer-vendor relationship

For vendors building AI agents, certification tells procurement teams, legal departments, and enterprise buyers that the system has been independently tested and gets re-tested on an ongoing basis, not that a policy document exists somewhere in a shared drive nobody opens. The insurance layer changes the conversation in a concrete way: a buyer using a certified, insured agent has actual coverage against the specific failure modes the standard tests for, and that reshapes how liability gets negotiated during procurement. As AI-specific questions creep further into security questionnaires, certification hands vendors a structured, already-documented answer about hallucination controls, prompt injection defenses, and unauthorized action prevention, instead of a scramble to explain it from scratch every single time a new deal comes up.

For organizations deploying or buying AI agents, AIUC-1 offers a real signal to check before handing an agent access to sensitive systems or customer data. Skipping that check, treating the badge as decoration rather than evidence, is the exact mistake this whole standard exists to prevent. The quarterly retesting cadence means a certificate issued this quarter reflects this quarter's threat landscape, not a stale picture from a year back. And with EU AI Act enforcement landing in August 2026, organizations that already built out their AIUC-1 evidence base show up with documentation already mapped to Articles 9, 11, 13, 14, 15, and 52, rather than starting that mapping exercise cold under a deadline.

Underneath all of it, the questions AIUC-1 is really answering (what can this agent access, what's it authorized to do, who's responsible when it does something nobody expected) are just extensions of identity, access, and operational risk questions that audit professionals have been asking for decades. The subject matter is new. The discipline behind it isn't, and mistaking the novelty of agentic AI for a need to reinvent audit practice from scratch is exactly the overcorrection AIUC-1 seems built to avoid.

For CPA firms and assurance practitioners, AIUC-1 runs on the same trust logic that SOC 1 and SOC 2 examinations have run on for years: independent, evidence-based proof that a system does what it claims to do. As AI agents start touching financial workflows directly, advisory firms with real assurance expertise are the ones positioned to help clients figure out where AIUC-1 actually fits inside a broader compliance picture. Treating it as one more badge to collect misses what makes it different in the first place, and that misreading is likely to be the most expensive mistake buyers make with this standard over the next few years.

Sources

  1. What Is AIUC-1? The Framework for Securing Agentic AI Systems
  2. What Is AIUC-1? The First Security Standard Built for AI Agents | Workstreet
  3. AIUC-1: Standard for AI Agents & AI Compliance Framework | K2 GRC Blog
  4. uipath.com

More in AIUC-1 and AI Assurance