Est.

AI Risk Assessment for Accounting and Finance Functions

Firms are deploying AI faster than governance structures can manage the resulting risks.

Features Editor · · 11 min read
Cover illustration for “AI Risk Assessment for Accounting and Finance Functions”
AIUC-1 and AI Assurance · September 15, 2026 · 11 min read · 2,583 words

AI is no longer a pilot project in accounting and finance. KPMG's 2024 survey found 72% of companies already using AI selectively in financial reporting, with adoption projected to hit 99% within three years, and Gartner reports 41% of internal audit teams already using or planning to use generative AI within the year. That speed is not the problem. Adoption has outrun the regulatory standards, documentation habits, and plain governance structures needed to manage it, and a vague warning about "AI risk" doesn't help anyone actually assess it. This piece breaks that risk into five categories, each with its own logic for evaluation: data quality, explainability, accountability, regulatory gaps, and third-party exposure, closing with a case study in how these risks stack up in a heavily regulated niche.

Data quality and model reliability as the foundation of every AI risk assessment

Every AI system in finance inherits the quality of whatever it's fed. Garbage in, garbage out, except at enterprise scale, where the garbage is spread across a chart of accounts, a dozen subsidiaries, and years of transaction history that may not even agree with itself.

Three things break first. Completeness: gaps in transaction records, missing counterparty fields, chart-of-accounts mappings that don't line up across entities that grew through acquisition instead of design. Integrity: unauthorized edits, data poisoning, or a model configured wrong by someone who didn't fully understand what it was supposed to do. SOC 2's Processing Integrity criterion addresses how system processing holds up under scrutiny, and it's a useful lens for catching hallucinations or strange model outputs, though stretching it that far is an interpretive reach, not something the standard explicitly demands, and the criterion is optional to begin with. Historical bias also shapes these models: one trained on five years of transaction data has absorbed five years of whatever accounting judgment calls, business cycles, and control weaknesses happened to be sitting there at the time, regardless of whether any of it still applies today.

So what does a finance leader actually ask? Not for the accuracy number in the vendor's pitch deck, but the number the tool produces against the organization's own data, in its own environment. Then how it behaves on edge cases matters: a novel transaction type, a merger, or a quarter disrupted by something nobody modeled for. Then whether anyone has a documented process for catching model drift before it shows up as an error in a filed statement.

The research gives reason for both optimism and caution, and the caution matters more than the headline number. A predictive AI model applied to audit lead identification hit 87% accuracy and explained close to 94% of the variance in loan disbursement amounts, according to a study published in Frontiers in Artificial Intelligence in January 2026. Strong numbers, but they came out of a controlled study, and controlled conditions rarely survive contact with a messy general ledger. Separately, Random Forest models applied to audit risk identification across Big Four audit data from 2020 to 2025 landed an F1-score of 0.9012, solid by most standards. But the feature importance analysis behind that score flagged factors including auditor workload and client ratings alongside financial metrics as key predictors. Which means the model's reliability rides partly on staffing decisions the finance team doesn't even control.

The real move here is mapping, and skipping it is the most common mistake a finance team makes. Every AI tool in the accounting stack needs a paper trail showing what data goes in, what comes out, and which control catches an error before it reaches the ledger or a financial statement. Without that map, "the model is accurate" is just a line from a sales deck that nobody can check.

Explainability and the "black box" problem in financial decision-making

Accuracy and explainability get treated as the same thing constantly, and that mix-up is where a lot of AI risk assessments go wrong. A neural network or ensemble model can spit out a correct answer with no auditable reasoning behind it, and in finance, that gap matters the moment the output touches a material judgment.

Under PCAOB standards, significant audit judgments are the responsibility of the engagement team, not a model. Where AI helps form that judgment, the workpaper file has to document the human review, not just paste in the AI's conclusion and call it done. Auditors face a practical expectation that they can explain, to a client or a regulator, why an algorithm flagged a particular transaction or item as high-risk. A score with no rationale attached is precisely the kind of documentation gap that creates exposure during an inspection.

Where does this actually show up day to day? Take journal-entry flagging: if AI marks a transaction as anomalous, can a controller or auditor point to the specific features that triggered the flag, or is it just a number with nothing behind it? Financial close raises the same issue: when AI-assisted reconciliation passes a balance, is there a record showing a human reviewed the logic, or did someone just glance at the output and move on? Disclosure drafting, too: if an LLM writes the first draft of MD&A language, what documents the human judgment applied before that language ever gets filed?

The assessment itself is easy to state and hard to execute. For each tool, check whether a finance professional can interpret its output without calling the vendor's support line. If they can't, find where in the workflow human review is supposed to happen, and what that review actually leaves behind as documentation.

Explainability is not a regulatory checkbox. It's an internal control question at its core: a control nobody can explain is a control nobody can test, and a control that can't be tested gives no one any real assurance.

Accountability and professional liability when AI participates in financial reporting

When AI makes an error in a financial statement, an audit opinion, or a tax filing, the liability lands on the CPA, not on the vendor whose model produced the mistake. US accountants operate under AICPA competence rules, SEC guidance, PCAOB audit standards, and a growing patchwork of state AI laws, and none of that shifts just because a third party built the tool.

That makes vendor reliance its own category of risk, one most firms underestimate. Audit firms using third-party AI tools need a real basis for concluding those tools fit the specific audit context they're used in. A SOC 2 report from the vendor does not by itself establish fitness for purpose. Fitness for purpose means checking the tool's methodology, its training data, and its scope against the actual engagement, not just confirming a compliance report exists somewhere in a shared folder.

One problem sits entirely unresolved right now: whether AI-enabled testing of 100% of journal entries satisfies PCAOB standards in place of traditional sampling. The scenario has been raised where inspectors could penalize a firm for using full-population AI testing precisely because no standard defines what counts as an acceptable AI-based audit procedure. Until PCAOB issues guidance on this, any firm running full-population AI testing is doing so without a defined benchmark to measure against, which is an uncomfortable place to sit during an inspection.

Finance leaders carry a parallel version of this, separate from what auditors face. Any AI tool helping prepare or analyze financial information is already part of a company's internal control over financial reporting, whether or not anyone bothered to classify it that way, and management owns those controls regardless. The readiness step is work that comes before the tool ever gets deployed: documenting how it was selected, how it's supervised, and what evidence shows someone actually checked its work.

So name a human owner for every AI tool touching the financial reporting chain, someone accountable for what it produces, and document the review process itself, not just how the tool happens to be configured.

The regulatory standards gap and what it requires organizations to do right now

No binding PCAOB or SEC standard governs AI in audits as of mid-2026. This carries real weight. It's the fact shaping every decision finance leaders make about AI deployment right now, and treating it as settled ground is the mistake to avoid.

Movement is happening, just not fast enough to close the gap yet. PCAOB's Technology Innovation Alliance Working Group finished its "Future State Deliverable" in May 2024, but the document wasn't released publicly until August 2025, a 15-month delay. As of mid-2026, none of the TIA's four strategic pillars have become binding standards. New PCAOB Chairman Demetrios (Jim) Logothetis issued RFC No. 2026-001 to help shape the 2026-30 strategic plan, with AI in financial reporting named as one focus area among several. That signals where the board's attention is heading. It is not a rule yet, and organizations planning around it as though it were are building on sand.

What's already binding tells a more useful story. AS 1105, covering audit evidence, is now fully in effect for calendar-year 2026 audits, and it requires auditors to evaluate the relevance and reliability of information obtained or processed using technology-based tools; when testing a company's controls over electronic information, that testing has to cover IT general controls relevant to it. QC 1000 takes effect December 15, 2026, pushed back a year from its original date, and it requires registered public accounting firms to build a quality control system that proactively identifies and manages risks to audit quality, including risks that emerging technologies introduce. Alongside QC 1000, standards AS 2901, AS 1215, AS 1220, AS 2101, and AS 2110 are all being updated, and audit committees need to understand how each change shifts their auditor's documentation obligations. Amended IT general controls standards kick in for fiscal years starting on or after December 15, 2026.

The risk sitting inside this gap cuts both ways. Over-reliance means assuming AI already satisfies a standard that doesn't exist yet. Under-documentation means assuming informal review is good enough, when inspectors, once standards do land, may expect a lot more on paper than what's currently sitting in the file. Treat the absence of a binding rule as a reason to document more, not less: map current AI usage against AS 1105 now, and build QC 1000-ready documentation of AI governance well before the December 2026 deadline.

Third-party AI providers and subprocessor risk in the financial services supply chain

Organizations using OpenAI, Anthropic, Google's AI services, or any other third-party AI provider have to evaluate that provider as a subservice organization under SOC 2. Auditors ask what financial data goes to that provider, how it's protected once it's there, and whether the provider holds its own SOC 2 certification.

A company's own SOC 2 report attests to how it uses those large language models: its vendor risk assessment, its data-retention and training opt-out settings, its review of subprocessors' SOC 2 and ISO 27001 reports. Organizations are increasingly expected to produce both their own SOC 2 and evidence that the LLM providers underneath them are independently audited.

SOC 2 does not cover everything, and treating it as though it does is where a lot of vendor risk assessments quietly fail. It covers the systems and processes wrapped around AI use. It does not certify model governance, bias handling, or explainability, none of which SOC 2 was ever built to measure. That's why many organizations layer ISO/IEC 42001, the AI management system standard published in December 2023, or the NIST AI Risk Management Framework's Generative AI Profile (NIST-AI-600-1) on top of SOC 2, specifically to close the gap SOC 2 leaves open. Some go further and test against Annex A of ISO 42001, 38 controls in total, incorporated into their SOC 2 report to show things like responsible-use controls actually exist.

COSO's Internal Control, Integrated Framework is the connective tissue underneath all of this. The AICPA's Trust Services Criteria are grounded in COSO, and COSO's principles give organizations a foundation for building AI into the SOC 2 control environment: governance, risk assessment, privacy, data protection, bias and fairness, and accountability for decisions the AI helps make.

Two control families matter specifically for AI agents. CC8.1 requires documented change management processes, covering how changes are initiated, assessed, approved, and tracked through to production, with mechanisms to address failures. CC7.2 and CC7.3 address continuous monitoring, covering the collection and analysis of security event data across critical systems. Best practice, though not something these criteria explicitly mandate, is a centralized logging repository with immutable storage held for at least a year.

Inventory every third-party AI provider touching financial data. Get their SOC 2 and ISO 27001 reports and actually read them. Document data-sharing agreements, training opt-out settings, and retention policies, and confirm subprocessor review shows up in the organization's own SOC 2 controls, not just in a vendor contract sitting in a drawer somewhere.

How AI risk compounds in specialized regulatory environments: the HUD 232 case

Every risk category above compounds when the regulatory environment gets more specific, and HUD's Section 232 program, which insures mortgages for nursing homes and other healthcare facilities, makes a useful case study for how that compounding actually plays out.

As of June 2024, 167 of 3,670 HUD-insured Section 232 borrowers, nearly 5%, had defaulted, carrying an unpaid principal balance north of $1.1 billion, according to a federal oversight watchdog. A separate HUD OIG review looked at four portfolios covering 70 properties, 84 loans, and a collective unpaid balance over $410.6 million, and found HUD didn't always act on risks sitting right there in borrowers' own audited financial statements.

Why not? Capacity, mostly. Staff at the HUD unit overseeing these loans didn't have the time to run the detailed analysis needed to catch unauthorized distributions, meaning cases where borrowers pulled funds out improperly. The staffing math explains why: ORCF headcount dropped from 61 employees in October 2024 to 38 by May 2025, while the average number of properties managed per staff member jumped from a range of 60-80 up to 165-200. That's a workload more than doubling in under a year.

That collapse in oversight capacity is, oddly enough, the strongest argument for AI-assisted portfolio monitoring in this entire piece. When human review capacity falls off a cliff like that, AI anomaly detection and automated flagging of covenant breaches or unauthorized distributions stops being a nice-to-have efficiency upgrade. It becomes the thing standing between a lender and a $1.1 billion default pile getting worse.

But how does this connect back to everything covered above? HUD 232 compliance runs on precise, regulation-defined financial reporting, and an AI tool that performs well in general-purpose bookkeeping can still misclassify transactions against HUD's specific chart-of-accounts requirements: the same data-quality risk from the opening section, just with sharper teeth. Catching unauthorized distributions requires judgment calls about operator intent, related-party relationships, and cash flow patterns that don't reduce cleanly to a score, which is exactly where explainability stops being optional and becomes a requirement for regulatory defensibility. And any AI tool touching the financial reporting chain for a HUD-insured property becomes part of the internal control environment an auditor has to evaluate under the amended PCAOB standards discussed earlier.

The assessment move for HUD-insured borrowers and their auditors follows the same five-part logic laid out across this piece, just at higher stakes: map the data, document the explainability trail, name the accountable human, track the tool against AS 1105 and QC 1000, and vet every subprocessor touching the numbers. The categories don't change. The tolerance for getting any one of them wrong does.

Sources

  1. Frontiers | Enhancing audit quality and reducing costs: the impact of AI in banking and financial services
  2. AI in Financial Reporting Audit Risk: The 2026 Compliance Map for CFOs and Audit Committees
  3. HUD Did Not Always Address Risks Reported in Borrowers’ Audited Financial Statements for Section 232 Residential Care Facility Portfolios | Office of Inspector General, Department of Housing and Urban Development
  4. hudoig.gov

More in AIUC-1 and AI Assurance