Est.

AI Model Documentation Requirements for Auditors

Auditors must document how AI assists their work, not just that it's used.

Staff Writer · · 13 min read
Cover illustration for “AI Model Documentation Requirements for Auditors”
AIUC-1 and AI Assurance · September 15, 2026 · 13 min read · 2,895 words

AI now sits inside audit workflows for financial statements, SOC examinations, mortgage banking, skilled nursing, and HUD-regulated housing. It's not a pilot program tucked away in some sandbox. It's already doing work that used to belong entirely to a human reviewer, and the documentation duties that come with it are no longer theoretical. This piece maps what those duties actually look like across the areas where AI is doing the most consequential work, not what regulators might eventually decide they should look like.

What existing auditing standards already require when AI touches the work

Start with the standards that already exist, because that's where most of the confusion lives. PCAOB's AS 1105, the standard governing audit evidence, requires auditors to check the relevance and reliability of information obtained or processed using technology-based tools. Where AI assists, the IT general controls tied to that information have to be tested. That's not a new rule invented for AI. It's an old rule with a new kind of tool to apply it to.

AS 1215, the audit documentation standard, got amended paragraphs (.09 and.11) approved by the SEC, and those amendments take effect December 15, 2026. What they demand, in plain terms: the file has to show how technology-based tools were actually used in the engagement. Not just that a tool sat somewhere in the process, but how it touched the work.

Here's where this piece takes a hard stand. Significant audit judgments cannot be handed off to AI, full stop. Any firm treating an AI output as a finished judgment rather than an input to one is already out of compliance, whether an inspector has caught it yet or not. Where AI assists in forming a judgment, the file needs to record the human review process that led to the final conclusion. The AI can suggest, flag, summarize, or draft. It cannot decide, and the file has to prove that line held.

PCAOB hasn't issued a formal safe harbor confirming that existing standards even permit AI use the way firms are currently deploying it. Firms are working off an assumption, which is a strange place for an entire profession to sit while moving this fast. One PCAOB board member raised a scenario worth sitting with: what happens when an AI tool lets a firm test 100% of journal entries instead of a sample, and an inspector penalizes that firm anyway, because no standard clearly defines what an acceptable AI-based audit looks like? Testing everything sounds like more rigor, not less. But without a standard drawing the boundary, more can look like a violation just as easily as it looks like diligence.

PCAOB tried to close some of that gap with Staff Guidance issued in October 2025, offering specific examples on evaluating electronic information reliability under AS 1105. Useful, but it's guidance, not a binding safe harbor. The gap between what firms need and what they've been given is still real.

The AICPA's framework for private company audits splits the obligation into several areas covering both firm-wide governance and engagement-level execution, including considerations around tool selection, vendor security, model constraints, and Confidential Client Information rules. Both the firm-wide piece and the engagement-specific piece need documentation. Treating one as a stand-in for the other misses the point of the framework entirely.

The IIA's AI Auditing Framework, updated in 2023 to line up with the NIST AI Risk Management Framework and the reality of large language models, organizes the same duties across governance, management, and internal audit functions. It covers strategic alignment, ethical risk, data governance, technical resourcing, third-party controls, and ongoing monitoring. That's a wider net than most firms are used to casting.

A December 2025 study in Advances in Accounting, surveying 89 IT auditors, found that the top auditability measures for AI-driven processes were log maintenance, standards compliance, and data governance processes. Sit with what that tells you: the old IT auditability toolkit doesn't go away. It just gets a second layer bolted on top, built around model explainability and data governance specifically.

The Model Bill of Materials: the documentation artifact auditors are increasingly expected to produce or demand

A deployment config file does not satisfy any of this. Neither does a fine-tuning notebook sitting in someone's repository. What auditors and regulators are converging on instead is a fuller artifact called the Model Bill of Materials, or MBOM. Why is that convergence happening now, rather than five years ago?

The EU AI Act, the first binding AI regulation anywhere, entered into force in August 2024. Articles 11 and 13 require documentation of AI system characteristics, including information about the model's design, data, intended use, and human oversight mechanisms. GPAI obligations under the Act took effect August 2, 2025. High-risk enforcement was originally set for August 2026, but the Digital Omnibus on AI, enacted July 27, 2026, pushed Annex III high-risk obligations out to December 2, 2027. That's a fixed deferral, not one that depends on harmonized standards showing up first.

The reach extends further than firms based in the region writing the rules. Any AI system whose output gets used within the covered region falls under it, so an audit firm based outside that region but serving clients with operations there faces the same documentation duties as a firm headquartered inside it. Distance from the regulator's home city buys no exemption.

On the technical side, CycloneDX ML-BOM version 1.7 gives the format a concrete structure. NIST's AI RMF and other national AI governance frameworks line up on the same underlying principles, even where the paperwork looks different. That convergence means something: three separate regulatory traditions landing on roughly the same list of fields to capture suggests the list itself reflects something close to consensus, not just regulatory preference.

An appliedAI study of 106 systems, done in March 2023, found that 40% of enterprise AI systems couldn't be clearly sorted as high-risk or low-risk. Sit with that number for a second. Without MBOM-level documentation, risk classification stops being a judgment call and turns into guesswork, and guesswork dressed up as classification is itself a documentation failure.

The core fields an MBOM needs to cover, across every practice area this piece touches: model identity (name, version, architecture type, whether it's a narrow tool or a fine-tuned LLM), training data provenance including known gaps or biases, intended use and explicit use restrictions, pre-deployment bias and fairness checks, human oversight mechanisms and escalation paths, a change history tracking retraining events and version control, and a list of third-party and vendor dependencies. ISACA's Advanced in AI Audit credential and its AI Resource Center both push auditors toward treating this as a living record, updated at each retraining or deployment change, not a snapshot filed away and forgotten.

Financial statement audits: what the file must show when AI assists the engagement team

2026 sits in an odd spot: a compliance year and an ambiguity year at once. AS 1105 amendments are already in effect. AS 1215 amendments land December 15, 2026. But no binding AI-specific standard exists yet telling firms exactly how to satisfy either one when AI is doing the work.

So what does the file actually need to show? How the AI tools touching financial information were picked, and the reasoning behind that pick. Vendor security due diligence completed before any client data got fed into the tool. Model constraints documented specifically to block unauthorized disclosure under Confidential Client Information rules. IT general controls tested against whatever electronic information the AI processed or produced. Human review steps for every significant judgment, proving the auditor made the call rather than rubber-stamping an output. And supervision evidence: who reviewed the AI's work, when, and what they concluded from it.

The 100% testing question raised earlier hasn't been settled. Where AI makes it possible to test an entire population of journal entries instead of a sample, an inspector might still challenge whether that satisfies PCAOB standards as written. The safer move right now, and the one worth taking over waiting for clarity that hasn't come, is documenting the reasoning behind the approach and the human review applied to whatever the AI flags. Assuming that more coverage automatically buys less scrutiny is a bet firms shouldn't make yet.

One point that gets missed: AI tools helping prepare or analyze financial information are already part of internal control over financial reporting. Their governance needs documenting on that basis, inside the control environment, not treated as some outside convenience bolted onto it.

PCAOB's request for public comment, 2026-001, issued under Chairman Logothetis, is currently gathering input to shape standards for 2026 through 2030. Firms documenting their AI practices carefully now aren't just protecting the current engagement. They're building the record that puts them ahead once binding standards actually land.

SOC 1 and SOC 2 examinations: why AI agents break traditional control assumptions and what replaces them

AI agents that write code on their own or make decisions without a human clicking approve break the assumptions built into traditional SOC 2 change management and access controls. The criteria themselves haven't changed. Applying them to an environment where an agent can act on its own needs a documentation approach that didn't exist when those criteria were written.

Take logical access, covered under CC6.1 and CC6.2. In an AI environment, those controls now have to reach model repositories, training datasets, and inference APIs. Documentation needs to show role definitions, least-privilege access enforcement, periodic access reviews, and immutable logging of both inputs and outputs. Skip any one of those and the control has a gap that a traditional access review wouldn't catch.

Change management, under CC8.1, gets more demanding still. Organizations deploying AI agents need a change request log recording who started each model update or retraining event, why, what risk assessment got performed, and whose signature approved it. They need automated version control tracking code revisions, with deployment locked to environments that already passed security testing. They need rollback mechanisms with documented triggers, ready before they're needed rather than improvised after something breaks. And they need periodic checks confirming that only approved changes actually reached production.

Monitoring controls under CC7.2 and CC7.3, covering continuous collection and analysis of security event data, now have to fold in AI systems as part of that collection. Anomaly detection coverage for those systems needs its own documentation, not an assumption that it exists just because it exists everywhere else.

Processing integrity criteria push straight into the model's development phase: approved training datasets, the reasoning behind feature selection, architecture decisions, validation mechanisms, error detection procedures, correction workflows. All of it needs a paper trail.

Here's a finding worth sitting with. CyLab research found that manipulating as little as 0.1% of a model's pre-training dataset is enough to launch an effective data poisoning attack. Think about what that means in practice: an attacker with nothing more than read-only access, someone who can submit documents into a training pipeline, could compromise the model without ever touching a write permission. Traditional integrity checks, built for access control, don't catch semantic manipulation buried inside a submitted document. So the auditor's job now carries a question that didn't used to need asking: do content integrity controls exist for training data, separate from and on top of access controls?

Evidence currency matters more here than almost anywhere else. AI policy documents are generally fine if reviewed within the past 12 months and signed by someone with real accountability. But risk assessments have to reflect the model as it exists right now, not as it existed at some earlier deployment. Auditors regularly find risk assessments that predate the live system by 18 months or more, and a stale assessment like that satisfies no framework's requirements, no matter how thorough it looked the day it got written. If the model got retrained or its deployment context shifted, the risk assessment needs updating to match.

HUD 232 and multifamily housing audits: applying general AI governance obligations where no specific rule yet exists

HUD's deadline structure doesn't bend for AI. Audited financial statements, the auditor's opinion, and a report on ownership and compliance are due no later than 90 days after fiscal year-end, no matter which tools got used to prepare them.

The stakes behind that deadline are climbing. As of June 2024, 167 of the 3,670 HUD-insured Section 232 borrowers, close to 5%, had defaulted on their mortgages. Rising defaults mean the relevant oversight office is watching financial reporting quality more closely, not less. Between October and December 2025, HUD OIG's single audit oversight activities included desk reviews of single audits submitted between April and June of that year, which reads as an active posture rather than a passive one.

Here's the gap worth naming directly: HUD has no explicit AI model documentation standard specific to Section 232 or multifamily programs as of mid-2026. Nothing in the regulatory text tells a firm exactly what to file when an AI tool touched the numbers.

An absent rule is not an absent obligation, though, and treating the silence as permission to skip documentation is the mistake worth calling out plainly. AI tools deployed in financial reporting or monitoring for HUD-regulated entities inherit documentation duties from two directions at once. First, from HUD's Consolidated Audit Guide, built to help independent auditors perform audits of profit-motivated program participants in HUD housing programs. Second, from OMB M-25-21 and whatever federal AI governance guidance applies more broadly.

In practice, that means naming the AI tools used in financial statement preparation, by name and version, not just by category. Recording the data sources feeding those tools. Documenting any model outputs that ended up in notes or compliance certifications. Capturing the human review steps applied before final submission. And confirming, explicitly, that financial reports contain reliable data per HUD's own compliance requirements, with the reliability check documented wherever AI contributed to that data.

The lesson here generalizes well past HUD. When a specialized program has no AI-specific rule of its own, auditors document AI use under whatever financial reporting standard already governs the program, layered with applicable federal AI governance guidance. The absence of a rule is not permission to skip the paperwork. It's an instruction to combine two existing rulebooks instead of waiting on a third one that hasn't arrived.

Skilled nursing facility audits: why AI touching Medicare billing codes creates a distinct documentation obligation

AI tools in skilled nursing facilities show up most often in two places: MDS coding and clinical documentation, where ambient scribes generate draft notes from a clinician's conversation with a patient. Both functions feed straight into Medicare billing codes under one payment model, which means both carry financial consequences the moment they're wrong.

Does it matter, from an auditor's standpoint, whether a coding error came from a tired nurse or a misfiring model? It doesn't, functionally. AI suggestions influence case-mix index determinations the same way human judgment does, so an AI-driven coding error looks exactly like a human coding error once it shows up in a billing record. That equivalence is precisely the exposure at stake, and it's why the audit trail matters more here than almost anywhere else in this piece.

A 2025 OIG audit of Pinnacle Multicare Nursing and Rehabilitation Center found an estimated $31.2 million in overpayments. Auditors traced the overpayments to billing that happened when the medical record didn't support the reimbursement rate code assigned, when services went to people who didn't actually need skilled nursing care, and when documentation requirements simply weren't met. Where an AI tool had a hand in generating those records or suggesting those codes, its role becomes direct audit evidence, not a footnote.

CMS is running an active validation program now, starting fall 2025 and run by a designated contractor. Beginning in fiscal year 2026, a randomly selected group of SNFs is required to submit documentation validating a defined set of MDS assessment records. This isn't some future audit risk sitting on the horizon. It's a program running in real time, pulling real facilities into review right now.

And the accuracy numbers on the AI side aren't reassuring yet. Research has found meaningful error rates in ambient-scribe-generated draft notes under controlled testing conditions. Read that again: seven out of ten draft notes carried errors before a human ever touched them. Any facility treating human review of AI-generated notes as a formality rather than the actual safeguard is misreading what that number means. It's the only thing standing between a draft note and a billing code that doesn't hold up under scrutiny, and the file has to show that review actually happened, not just that a review step existed on paper.

What auditors need documented when AI touches SNF billing and clinical records: the identity and version of whatever AI scribe or coding-assist tool got used, the workflow showing the AI draft moving to a CDI specialist or clinician for review before it gets finalized, clear evidence that the human reviewer's judgment, not the raw AI output, is what the billing code rests on, a record of any instances where a clinician overrode an AI suggestion and why, and an honest comparison of vendor accuracy claims against the facility's own error rate monitoring. That last point matters more than it looks. A vendor's marketed accuracy rate and a facility's actual, observed error rate are two different numbers, and only one of them belongs in the audit file as evidence.

Sources

  1. 2025 Volume 16 An Auditors Guide to AI Models Considerations and Requirements
  2. Model Bill of Materials, EU AI Act, AI BOM, ML-BOM | Medium
  3. Bridging IT auditors and AI auditing: Understanding pathways to effective IT audits of AI-driven processes - ScienceDirect
  4. AI Governance Library

More in AIUC-1 and AI Assurance