Est.

AI Transparency and Explainability in Audit Contexts

Auditors need to defend AI decisions to regulators, not just understand them.

Staff Writer · · 16 min read
Cover illustration for “AI Transparency and Explainability in Audit Contexts”
AIUC-1 and AI Assurance · September 16, 2026 · 16 min read · 3,565 words

AI now operates inside audit workflows, not just around them, and that changes what "explainability" has to mean. When a machine learning model flags a journal entry as anomalous or classifies a lease under a specific accounting standard, the real question is whether the auditor who relied on it can defend that reliance to an audit committee, a PCAOB inspector, or an external reviewer who wasn't in the room when the output got generated. It's whether the auditor who relied on it can defend that reliance to an audit committee, a PCAOB inspector, or an external reviewer who wasn't in the room when the output got generated. That standard is narrower and, in some ways, harder than the one governing AI in lending or insurance, for reasons that matter before looking at what firms are actually doing about it.

Most of the public conversation about explainable AI in finance runs through credit decisions: loan denials, fraud flags, insurance pricing models that regulators want documented for fairness. That's a real and active regulatory space. But audit sits in a different position. The audience for an audit explanation is not a regulator checking whether a model treated a demographic group equitably. It's an audit committee, an external reviewer, or a PCAOB inspector checking whether a qualified professional exercised, and documented, appropriate judgment. Those are two different questions wearing the same word, "explainability," and conflating them leads firms toward the wrong kind of documentation.

This distinction predates AI by decades. AU-C 620 already requires auditors to evaluate the methods and work of any specialist they rely on. PCAOB AS 1201 puts the engagement partner on the hook for the engagement's performance, including supervising work so it actually supports the conclusions reached. Neither standard cares what tool produced the underlying work product. An analyst's spreadsheet, a specialist's model, a machine learning classifier: the auditor's obligation to understand and stand behind the result doesn't change based on what generated it. What AI does is make the gap between "I ran it through the tool" and defensible documentation bigger, and the stakes of that gap higher. It was always a documentation problem. It's just a louder one now.

What explainability means in an audit workflow (three layers)

Break explainability down and three distinct layers appear, each answering a different question, each failing in a different way.

Transparency comes first, and it's more basic than people assume: what data did the AI actually use, and can the output be traced back to a source document? The immediate task is understanding what data the AI actually used and whether the output can be traced back to a source document. It's about whether there's a connective thread between an invoice, a contract, a bank statement, and the number the AI produced. No thread, no starting point for anything else.

Interpretability sits above that. Can a senior accountant or auditor look at the AI's output and understand why it arrived there, well enough to exercise professional judgment over it rather than just accepting it? This is where a lot of AI tooling quietly fails. A confidence score is not an explanation. A classification label is not a reason. If the senior on the engagement can't articulate why the model flagged what it flagged, no professional judgment actually got applied, no matter what the workpaper says.

Traceability is the layer that matters most for audit specifically, because it's the one a PCAOB inspector will actually test. Can the path from input to output be reconstructed in a way that fits into the evidentiary structure of a workpaper? An auditor doesn't get to say "the model said so." The reconstruction has to hold up on its own, months or years after the fact, to someone who wasn't there when the analysis ran.

Apply this test to each layer separately, because firms tend to assume passing one means passing all three: for transparency, can the reviewer actually pull up the source document the AI drew from? For interpretability, can a senior auditor evaluate the reasoning behind the output, not just read the conclusion? For traceability, does the output slot into the workpaper structure a PCAOB inspector would expect to see?

Consider the scenario that plays out when external auditors test AI-generated financial data and ask the client's controller to explain how a number was derived. If the controller's answer is "the platform calculated it," that failure spans a minor documentation gap that must be cleaned up later. It's a failure across all three layers at once, and it's the kind of answer that turns a routine test into a finding.

It helps to place all this inside a bigger concept: auditability. Research on the subject frames auditability as the full set of procedural, technical, and organizational conditions, documentation, data traceability, log access, model attributes, that make outside scrutiny possible in the first place. Transparency and explainability are contributing factors to auditability rather than synonyms for it. They're contributing factors to it. An AI system can be perfectly transparent about its inputs and still fail to be auditable if there's no organizational process for anyone to actually go check.

The technical approaches auditors and their tools use to produce explanations

Two broad technical families produce explanations, and they trade off against each other in a way that auditors need to understand before choosing a tool.

Interpretable models, things like decision trees and rule-based systems, are transparent by construction. The logic is understandable on its own, with no separate explanation layer bolted on afterward. The tradeoff is predictive precision: these models often catch less nuance than more complex alternatives, which matters when the fraud pattern or misstatement being hunted for is subtle.

Post-hoc methods go the other direction. They let auditors use more powerful, more opaque models, then apply a separate explanation technique after the fact to generate something a human can read. That flexibility is valuable, but it introduces its own reliability question: is the explanation actually describing what the model did, or is it a plausible-sounding approximation?

A few of these techniques have moved from academic papers into active audit use. Research from Zhong and Goel, working through a fraud detection classification case, demonstrated three of them in practice. SHAP (SHapley Additive Explanations) quantifies which input features drove a specific result, so the auditor isn't just told what the model concluded but which factors carried the most weight in getting there. LIME works locally: it approximates the model's behavior in the neighborhood of one specific prediction, which is useful when the auditor cares about explaining a single flagged transaction rather than the model's behavior overall. Counterfactual explanations ask a different question entirely: what would have needed to change in the inputs to produce a different output? That's a way of mapping the boundaries of the model's reasoning, useful for understanding how sensitive a conclusion actually is.

A systematic review published in Frontiers in Artificial Intelligence found that explainable AI applications in audit and finance cluster in a handful of areas: fraud detection, credit risk assessment, regulatory compliance, financial statement analysis, and decision-support processes generally. That's a narrower footprint than the AI hype cycle might suggest, and it tracks with where the stakes of getting an explanation wrong are highest.

But how much can any of this actually be trusted? The limitations are real and documented. A BIS FSI Occasional Paper from September 2025 lays out the problem: existing explainability techniques suffer from inaccuracy, instability, and susceptibility to misleading explanations, and complex models often remain difficult to explain even with SHAP or LIME applied. An explanation method can produce a clean-looking chart of feature importance that doesn't actually hold up if the underlying data shifts slightly.

That same BIS paper raises a question regulators haven't settled yet: whether complex, high-performing AI models with limited explainability should be allowed to operate anyway, provided adequate safeguards exist around them. Nothing about that is resolved. But the fact that a body like BIS is willing to float it signals the conversation has moved past "explainability or nothing" and into something more nuanced about acceptable tradeoffs.

The PCAOB and IAASB requirements, and the gaps that remain

The PCAOB named AI an area of emphasis in its 2025 inspection priorities, and that wasn't a symbolic gesture. Its Technology Innovation Alliance Working Group spent roughly 18 months developing recommendations on how oversight programs should handle emerging technology, which suggests the regulator is trying to build durable infrastructure rather than issue a one-off warning.

Recent inspection findings give that emphasis teeth. Deficiencies have concentrated in journal-entry testing, specifically firms using AI-assisted anomaly detection without documenting how the algorithm was calibrated, what data trained or informed it, or how false positives got evaluated. Under AS 1105, that's a deficiency regardless of whether the AI's output happened to be right. Being right isn't the standard. Being able to show the work is.

A real gap exists between current practice and what firms largely believe, though. Generative AI use in public audits today is concentrated in lower-stakes work: drafting memos, summarizing policy language, research tasks. Firms largely believe existing standards don't block broader adoption. But the PCAOB hasn't formally confirmed that reading. Firms are operating on an assumption, not a guaranteed safe harbor, and that's a meaningfully different footing to build a practice on.

The PCAOB did offer a partial bridge in late 2025, issuing Staff Guidance with specific examples on evaluating electronic information reliability under AS 1105. Partial is the right word. It addresses one standard, leaving open the broader question of how AI fits into the audit standard set as a whole.

The more ambitious signal came in the PCAOB's Future State Deliverable, released publicly in 2025, laying out four strategic pillars: standardizing audit documentation through structured data so AI tools (and continuous auditing generally) can actually work with it, a framework for responsible AI use by auditors, a Regulatory Innovation Lab functioning as an audit-tech sandbox for piloting new standards, and a push on auditor tech literacy through data analytics and AI training. None of that is finished rulemaking. It's a roadmap.

Then came a leadership and strategic reset. Demetrios (Jim) Logothetis was appointed PCAOB Chairman in early 2026, and the board issued public RFC No. 2026-001 to gather stakeholder input as a first step toward its 2026-30 strategic plan, with updating auditing standards for AI in financial reporting as a central focus. That RFC is a starting gun, not a finish line, and firms watching for concrete standard-setting will need to wait through the comment and drafting process.

The IAASB has been moving on a parallel track. Its October 2024 Technology Position committed the board to a gap analysis of whether current standards adequately address AI, folding technology considerations into ongoing standard revisions instead of waiting for the technology to settle down first. That's a deliberate choice: the standard-setter would rather adjust as it goes than freeze activity until AI stabilizes, which, given how fast the tooling is changing, may be the only workable approach.

The most consequential near-term change isn't AI-specific at all. SAS No. 146 moves firms toward risk-based quality management for periods beginning on or after December 15, 2025, and it is alongside the PCAOB's enhanced confirmation standard and the IAASB's revised going-concern standard, all arriving in the 2025 to 2026 period. Firms adjusting their quality management systems to satisfy SAS 146 have a natural opening to fold AI governance controls into that same infrastructure, rather than building a separate compliance track later. The AICPA's Auditing Standards Board is watching a related AI project from the Assurance Services Executive Committee, with an eye toward whether additional guidance is needed to keep AI use consistent with audit quality.

Explainability obligations under SOX, the SEC, and the EU AI Act beyond the audit firm

None of this stays contained inside the audit firm. SOX Section 404 requires documentation of financial controls, and that includes AI-driven decisions that affect reported financials. The auditor's explainability obligation and the client's documentation obligation are really the same compliance problem viewed from two sides of the engagement.

The SEC's 2025 Examination Priorities address AI use by registered firms directly, with the Division reviewing whether registrants maintain adequate policies to monitor and supervise AI use, including for fraud prevention, anti-money laundering, and protection of client information. To be precise about what this doesn't say, there's no specific SEC mandate requiring audit trails for every AI-driven decision, and no rule that algorithms themselves must comply with GAAP or IFRS. The focus instead falls on whether automated processes, in aggregate, adhere to those frameworks. That's a softer requirement than a hard audit-trail mandate, but it still puts the burden of proof on the firm.

The EU AI Act reaches further out and arrives on a longer timeline, requiring institutions deploying high-risk AI systems to maintain comprehensive traceability documentation once the relevant provisions take effect, and decision logs, once the relevant provisions take effect (December 2027 for standalone high-risk systems under Annex III, August 2028 for AI embedded in regulated products under Annex I). This isn't confined to firms headquartered in the EU. Any firm with EU operations or EU counterparties needs to account for it.

What happens when none of this documentation exists? Switzerland's FINMA flagged the real-world version of the problem in 2024, noted in a BIS FSI paper: some AI model results simply can't be understood, explained, or reproduced after the fact, which makes it difficult for regulators to confirm whether a financial institution is meeting existing requirements, particularly in business areas the regulator considers critical. That's a regulator on record saying the explainability gap already interferes with its ability to do its job. It's a regulator on record saying the explainability gap already interferes with its ability to do its job.

Which points to an architectural lesson that appears repeatedly in this space: audit trails built after the fact and retrofitted onto a system that's already in production tend to have gaps. Missing prompt versions. Incomplete source attribution. Timestamps that don't line up cleanly. Those gaps become visible exactly when they're most costly, during the review that actually needs the log to hold together. Explainability has to be designed into the system from the start. Bolting it on later doesn't produce the same result, no matter how much documentation gets generated afterward.

Where SOC 2 stops and AI-specific assurance begins

SOC 2 gets invoked constantly in AI vendor conversations, so it matters to be precise about what the report actually covers. Per the AICPA, a SOC 2 report addresses controls at a service organization relevant to security, availability, processing integrity, confidentiality, or privacy. Security is mandatory in every SOC 2 engagement; the other four Trust Service Criteria apply based on scope. Together they describe how a system is designed, monitored, and controlled operationally.

For an AI company specifically, a SOC 2 audit tests whether the organization controls its machine learning lifecycle: data ingestion, training, deployment, monitoring. Model integrity considerations, including whether a model has been altered or degraded over time, surface as a risk area within SOC 2 engagements.

There's a hard limit, though, and it's one that gets glossed over in vendor sales conversations more often than it should. SOC 2 attests to controls at the organization running the system. It tells a buyer the company has sound operational security practices around its infrastructure. It says nothing about how the AI itself reasons, what it draws its conclusions from, or whether its outputs can be explained to a third party. There's no clause in SOC 2 that's unique to AI or machine learning at all. It applies to an AI company exactly the way it applies to any cloud provider or SaaS vendor handling customer data, full stop.

That gap has become increasingly difficult to ignore. Growing recognition that AI governance controls carry real weight has pushed SOC 2 practice beyond its traditional periodic-checkpoint model, toward something closer to continuous, intelligent monitoring. That's a shift in posture more than a rewrite of the framework itself, and the frameworks actually built to close the AI-specific gap sit outside SOC 2's scope.

ISO/IEC 42001, the AI management system standard published in December 2023, covers model governance, bias, and explainability in ways SOC 2 was never designed to touch. NIST's AI RMF Generative AI Profile (NIST-AI-600-1, published July 2024) covers similar territory from a different angle. Because the two frameworks share overlapping controls, most AI companies pursuing rigorous assurance end up auditing against ISO/IEC 42001 and SOC 2 together rather than picking one. Further out, AIUC-1 is emerging as an AI-specific assurance framework aimed directly at how AI systems behave, reason, and can be examined, which is precisely the territory SOC 2 leaves open.

Explainability demands in specialized audit contexts (HUD, skilled nursing, and mortgage banking)

Abstract standards are one thing. Specialized regulatory environments make the stakes concrete fast, and HUD-insured multifamily and healthcare properties are a good place to see it.

Audited financial statements are required for profit-motivated multifamily projects with annual expenditures, or HUD-insured or HUD-guaranteed loan balances, of $500,000 or more. Non-profit projects face a different threshold: $1,000,000 or more in annual expenditures (not loan balances) for fiscal years beginning on or after October 1, 2024, with a $750,000 threshold for fiscal years beginning before that date. The HUD Consolidated Audit Guide lays out the audit procedures both for-profit and nonprofit entities must follow, and those procedures apply whether or not AI touched any part of the underlying analysis.

The default numbers give this weight beyond paperwork. HUD's OIG found that 167 of 3,670 HUD-insured Section 232 borrowers, nearly 5 percent, had defaulted on their mortgages. A default rate at that level intensifies ORCF oversight scrutiny across the whole portfolio. As a result, auditor documentation on these engagements carries more consequence than it might have five years ago. Section 232 engagements cover FHA-insured healthcare mortgages and 232/223(f) loans, including post-closing reporting for skilled nursing and assisted living operators. If an AI tool influences how occupancy revenue gets classified in that reporting, the auditor needs a traceable path back to source data, because the regulator reviewing that file already has reason to look hard.

HUD's Fair Housing Act guidance from May 2, 2024 adds another layer, specifically around AI-assisted tenant screening and advertising in the multifamily portfolio. That guidance creates real transparency exposure for AI screening decisions that aren't documented, and while it doesn't explicitly extend to HUD 232 healthcare facilities, it covers the overlapping multifamily segment closely enough that auditors of these portfolios need to understand whether the screening tools in use are governed and documented at all.

Skilled nursing accounting compounds the problem further, sitting at the intersection of HUD 232 post-closing requirements, Medicare and Medicaid cost reporting, and state licensure compliance simultaneously. An AI tool applied anywhere in that workflow has to be explainable across multiple regulatory audiences at once, not just one. Mortgage banking carries a parallel version of the same risk: investor reporting, agency delivery, and rep-and-warranty exposure all depend on AI-assisted underwriting or quality-control outputs being traceable back to source documents. Without that traceability, the exposure reaches beyond an audit deficiency. It's rep-and-warranty risk that can follow a loan for years.

The thread connecting all three contexts is the same: an auditor working these engagements has to defend conclusions to a regulator who understands that industry's specific risk profile in detail. A generic AI output with no traceable path back to source data doesn't hold up under that kind of scrutiny, no matter how accurate it turns out to be.

What an explainability-ready audit practice looks like in operation

Pull the threads from PCAOB inspection findings, IAASB positioning, SOC 2's limits, and the sector-specific cases together, and one architectural principle produces all of it: explainability has to be built into an AI-assisted workflow from the beginning. Retrofitted audit trails create exactly the gaps examiners find, missing prompt versions, incomplete source attribution, timestamps that don't line up, and those gaps become visible at the worst possible moment, when an examiner is actually trying to reconstruct what happened.

One condition matters more than the others in practice, and it's the one PCAOB inspection reports have already flagged directly. For any AI tool used in substantive testing, the engagement file needs to record how the algorithm was calibrated, what data fed into it, and how its outputs were evaluated once produced. That's the specific gap inspectors have cited in journal-entry testing, already a known point of failure rather than a theoretical one. It's the specific gap inspectors have cited in journal-entry testing. It's already a known point of failure rather than a theoretical one.

None of this is fully settled yet, and that deserves sitting with rather than smoothing over. The PCAOB's 2026-30 strategic plan is still being shaped through stakeholder input. The IAASB's gap analysis is ongoing. The BIS paper's open question, whether limited explainability might be acceptable alongside the right safeguards, hasn't been answered by any regulator with authority to answer it. What is settled is the underlying logic that predates all of this AI-specific activity: a qualified professional has to be able to show, on paper, why a conclusion was reached. AI just makes the absence of an answer a great deal harder to hide under that same obligation. It just makes the absence of an answer a great deal harder to hide.

Sources

  1. Transparent AI in Auditing through Explainable AI | Current Issues in Auditing | American Accounting Association
  2. Managing explanations: how regulators can address AI explainability | Bank for International Settlements
  3. Explainable artificial intelligence in accounting and financial auditing: a systematic review - PMC
  4. Frontiers | Explainable artificial intelligence in accounting and financial auditing: a systematic review
  5. calcpa.org
  6. researchgate.net
  7. lbmc.com
  8. Explainable AI in Finance: Addressing the Needs of Diverse Stakeholders

More in AIUC-1 and AI Assurance