Est.

AI Assurance vs Traditional IT Audit

Traditional audit misses AI drift, bias, and opacity that new standards must address.

Staff Writer · · 11 min read
Cover illustration for “AI Assurance vs Traditional IT Audit”
AIUC-1 and AI Assurance · September 15, 2026 · 11 min read · 2,571 words

Pick a sample. Test whether the controls in that sample existed and operated the way they were supposed to during the period under review. That method has governed IT audit for decades, under frameworks like COBIT 2019 and PCAOB AS 2201 for audits of internal control over financial reporting.

AS 2201 lays out the sequence: understand the internal control environment, assess where a material weakness could occur, then test whether controls were designed properly and operated effectively. It's a bounded engagement with a clear finish line, and that boundedness is the point, not a limitation someone forgot to fix.

This approach catches a lot. Access control failures, missing segregation of duties, undocumented changes, policy violations that leave a paper trail: all of it surfaces under this kind of testing. What it misses is misbehavior that emerges from a system whose internal logic shifts as new data comes in. The sampling constraint made sense when full-population testing wasn't computationally possible, and for financial controls that stay put between review periods, it still makes sense. A traditional IT auditor checks whether a control was in place. Whether a model's outputs are fair, accurate, or drifting over time sits outside that job description entirely, and no amount of rigor applied to the wrong question fixes that.

Where traditional IT audit runs out of runway with AI systems

AI systems, especially deep learning models, don't behave like the systems traditional IT audit was built around. ISACA's 2025 AI Pulse Poll and related blog content, published alongside its Advanced in AI Audit credential, say this directly: these systems carry a kind of complexity that sits behind logic nobody can fully trace.

A traditional audit can confirm a model got deployed. It cannot tell anyone what that model does with a specific input, or whether its behavior holds steady across different demographic groups or unusual edge cases. That's the opacity problem, and no amount of access-control testing solves it.

Then there's drift. A model that looked fine during a point-in-time review can quietly degrade as the world it operates in shifts underneath it. Nothing in a traditional control framework catches that, because the framework assumes the thing being tested stays roughly the same between reviews.

Sampling breaks down here for a specific, almost arithmetic reason. A small sample of model outputs is never going to surface a bias affecting a narrow slice of transactions concentrated in one subpopulation. That's not carelessness on the auditor's part, the math simply doesn't clear at that scale. Add algorithmic bias, adversarial attacks aimed specifically at model inputs, and data quality that degrades over time, and the result is a set of risks a controls checklist cannot see by design. Regulators are naming these AI-specific risks faster than audit standards have caught up to address them, which is exactly backwards from how assurance is supposed to work.

What AI-enabled auditing techniques actually change about evidence collection and controls testing

AI is not just the subject of audit anymore. It's also becoming a tool auditors use, and that dual role is where a lot of confusion creeps in, because using AI well as an auditing tool tells you nothing about whether the AI being audited is trustworthy.

Full-population testing replaces sampling in some contexts now. AI tools can scan entire datasets close to real time, flagging transactions or configurations that look off. ISACA's Now Blog discussed in May 2025 how continuous monitoring can replace what used to be a periodic snapshot with something closer to a live feed.

Process mining is part of the same shift. Instead of checking whether a control checkbox got ticked, AI can trace an entire workflow: how long a transaction sat at each step, where it got diverted, what exceptions came up along the way. Natural language processing does something similar for policy compliance, cross-referencing internal policy documents against regulatory text automatically and surfacing gaps that used to take a person hours of manual reading to find.

The numbers are worth sitting with. A structured narrative review covering 2020 through 2025 found AI-assisted techniques improving fraud detection speed by 40 to 50 percent, alongside a 30 to 40 percent cut in audit cycle time. Real gains, and worth taking seriously.

But look closely at what these gains actually measure. They make auditors faster and more thorough at traditional audit work. They don't, on their own, tell anyone whether an AI system behaves the way it claims to. Using AI to audit and holding assurance over an AI system are two separate questions, and the industry keeps answering the second one as though the first one had already covered it. It hasn't.

Why SOC 2, the closest existing framework for technology service providers, does not close the gap

SOC 2 attests to controls at the service organization level: security practices, availability, confidentiality, all measured against the AICPA's Trust Services Criteria. It's a well-established standard, and for a long stretch of the software industry's growth it's been the report buyers ask for before signing a contract.

Here's the mistake worth naming: treating a clean SOC 2 report as evidence of trustworthy AI. The point is worth making directly. A company can hold a spotless SOC 2 report and still run a model that produces biased credit decisions, hallucinates inside a financial workflow, or drifts silently once the examination period ends. The report and the model's actual behavior answer different questions, and familiarity with one doesn't substitute for the other, no matter how thorough the SOC 2 engagement was.

SOC 2's October 2022 guide update did modernize its points of focus to account for new and emerging technologies. That helped, somewhat. But the Trust Services Criteria still evaluate the organization's controls around a system, not what the system itself outputs. SOC 2+ examinations, layering in HITRUST, NIST CSF, or HIPAA mapping, extend what gets covered but stop short of behavioral assurance for the model itself.

Scale makes the gap easy to see. Large cloud providers' SOC reports span enormous footprints of organizational controls under examination. That's an enormous footprint of organizational controls under examination, and AI-specific behavioral assurance is nowhere in it. SOC 2 leaves unanswered whether a given AI system does what the vendor says it does, reliably, across every population it touches in production. Execution alone cannot close that gap. It's a gap in what the standard was ever built to ask.

What AIUC-1 is designed to answer that prior frameworks leave open

AIUC-1 is a newer standard aimed squarely at that gap. It evaluates whether AI-driven systems behave as intended, focusing on controls, accountability, and oversight mechanisms designed to evaluate whether AI-driven systems behave as intended.

What matters behaviorally is whether a system performs consistently, accurately, and fairly across the conditions it actually meets in production, not just the conditions it was tested against once, in a lab, before launch. That distinction sounds small until you consider how much can shift between a pre-launch test set and six months of live traffic feeding the same model.

It shows up in the evidence AIUC-1 engagements demand. Instead of confirming a model got approved for deployment, these engagements demand behavioral evidence rather than purely documentary evidence. Documentary evidence, the kind traditional audit runs on, gives way to behavioral evidence. That shift in evidentiary standard is the whole point of the framework, not a footnote to it.

None of this replaces conventional IT audit. It sits next to it. An organization running AI inside a financial workflow may genuinely need both a SOC examination for its organizational controls and an AIUC-1 engagement for the model's behavior, because the two engagements answer separate questions that both matter. Choosing one over the other, on the assumption that either one covers the other's ground, is where organizations get exposed.

Standard-setters are already leaning this direction. The PCAOB's 2026 request for comment (RFC No. 2026-001) and its Technology Innovation Alliance Working Group's Future State Deliverable, released publicly in September 2025, point at standardized AI documentation, responsible-use guidance, a regulatory innovation lab, and better auditor tech literacy. Standard-setters more broadly are pushing the same direction, asking firms to evaluate AI at both the engagement level and the firm governance level. Model behavior is turning into an audit-relevant question in its own right, not something IT quietly handles on the side.

The skills gap that makes AI assurance harder to staff than traditional IT audit

Auditing an AI system properly requires knowing something about the models themselves: data science, statistics, the math underneath the outputs. ISACA's 2025 research says this outright, and it's worth taking at face value, because it points at a staffing problem the industry hasn't solved yet.

Seventy percent of audit and assurance professionals who answered ISACA's 2025 AI Pulse Poll said they need to build AI skills within the next year just to stay competitive in their careers, let alone advance. That's a striking number for a profession that already runs on continuing education requirements.

Adoption inside internal audit functions hasn't kept pace, and the mismatch is stark once the two figures sit side by side. An estimated 55% of businesses are rolling out AI broadly, according to CrossCountry Consulting, while only 2 to 4% of internal audit teams have made any AI progress at all. Separately, 84% of surveyed professionals were found to cite workforce adaptation problems tied to weak AI training programs. Demand for auditors skilled in AI is rising faster than the supply of them, and that gap is exactly where firms overstate what they can actually deliver.

ISACA's answer is the Advanced in AI Audit credential, built around three domains: AI Governance and Risk, AI Operations, and AI Auditing Tools and Techniques. Most firms marketing "AI audit" services right now cannot actually back it up. A general cybersecurity background dressed up for a new market is not the same thing as lifecycle knowledge, and a firm that cannot explain how it tests for bias across subgroups, or how it distinguishes drift from noise, is selling a label rather than a capability. What buyers should look for is demonstrated command of the full AI lifecycle: data management, algorithm development, change management. Anything short of that is a credential mismatch waiting to surface at the worst possible time.

How organizations with AI in financial workflows should think about which engagement they actually need

Start with one question: is the AI system making, or materially shaping, decisions that touch financial statements, regulatory compliance, or third-party trust? If yes, a traditional IT audit of the surrounding controls is necessary but nowhere near sufficient on its own. Behavioral assurance of the model itself has to sit alongside it. An organization that stops at the traditional audit has only answered half the question it needed to ask, and treating that half-answer as complete is the exact error this piece keeps circling back to.

Picture three layers stacked on top of each other for an AI-enabled service organization. A financial statement audit covers the numbers. A SOC 1 or SOC 2 examination covers organizational controls relevant to user entities or trust criteria. An AIUC-1 engagement covers whether controls, oversight, and accountability mechanisms around the AI system are operating as intended. Skip any one of the three and there's a hole in the picture, not a redundancy removed. Treating any single engagement as "enough" misreads what each one was built to do.

Some industries feel this more urgently than others right now. Mortgage banking, where AI assists underwriting decisions. Skilled nursing facility billing, where AI drives coding and documentation. Multifamily housing compliance, where automated systems report against HUD requirements. Regulators are already paying close attention in each of these sectors, and a model error there carries financial and legal consequences that are anything but abstract.

Consider HUD's footprint alone: more than 26,000 Multifamily Housing and Office of Residential Care Facilities participants file annual electronic financial data with the agency. As AI moves into those reporting workflows, whether the model behaves correctly stops being a technology question and becomes a compliance question. Skilled nursing carries its own version of the same problem. The OIG's November 2024 Nursing Facility Industry Specific Compliance Program Guidance flagged four key risk areas, including Medicare and Medicaid billing requirements, exactly where AI-assisted billing tools are spreading fastest and where quiet model drift can turn into fraud exposure or overpayment liability.

CrossCountry Consulting's analysis frames this as part of a bigger shift for internal audit generally: the function is moving toward a strategic advisory role, one piece of which is helping the organization figure out whether its AI systems are being assured at the right level in the first place.

What a rigorous AI assurance engagement looks like in practice, and what it does not cover

Good AI assurance work leaves behind a specific kind of evidence: population-level output testing, bias analysis across the subgroups that matter for the use case, model drift monitoring across the full coverage period rather than a single snapshot, and documentation tracing where the training data came from and how changes to the model got managed along the way.

Explainability sits at the center of this. ISACA's Now Blog, in a November 2025 post, argued that IS auditors should favor AI models where outcomes can actually be traced back to their inputs. Black-box deep learning presents a specific problem here. If the internal logic can't be inspected, testing behavior at the output level becomes the only option left, not a preferred method chosen for convenience.

What the engagement does not do matters just as much, and here's where the piece has to be blunt: it does not replace a financial statement audit, it does not stand in for a SOC examination of organizational controls, and it does not certify that the model will never fail for the rest of its operational life. What it provides is reasonable assurance that the model behaved as represented during the specific period examined. That's a narrower and more honest claim than "certified safe." Any firm selling the latter is overselling what the methodology can support, full stop.

Some constraints don't disappear just because the assurance methodology improves. Data quality limits what any test can reveal. Bias testing is only as good as the representativeness of the test data itself, and biased test data produces a false clean bill of health, arguably worse than no testing at all because it manufactures confidence where none is warranted. Cybersecurity risks specific to model inputs, adversarial manipulation among them, remain live. And PCAOB and AICPA standards are still catching up to where the technology already sits. The PCAOB's own idea of a regulatory innovation lab, a sandbox for piloting standards before they get locked into formal rule-making, is itself an admission that the methodology has to keep evolving alongside the systems it examines.

For an organization choosing who does this work, the practical filter comes down to a narrow intersection: deep knowledge of how AI systems actually work, fluency in attestation standards like SOC and AIUC-1, and familiarity with the specific regulatory terrain, HUD, OIG, PCAOB, AICPA, that governs the industry in question. Very few firms sit at all three points at once, and that scarcity is the real story here, not a footnote to it. The right question for a buyer to ask is what a firm actually delivers when it offers "AI audit" services. It's whether the people doing the work actually hold all three.

Sources

  1. ISACA Now Blog 2025 The Evolving Connection Points Between AI and Audit
  2. ISACA Now Blog 2025 Embracing AI in Information Systems Audit A New Era of Assurance
  3. ISACA Now Blog 2025 Why IT Auditors Need AAIA Addressing AI Challenges in Audit
  4. Future of Internal Audit in 2025: AI, Cyber & Regulatory Risk
  5. schellman.com

More in AIUC-1 and AI Assurance