Est.

Internal Controls Testing in External Audits

Weak controls expand audit scope and cost; strong ones compress it.

Senior Writer · · 14 min read
Cover illustration for “Internal Controls Testing in External Audits”
Financial Statement Audit · July 28, 2026 · 14 min read · 3,104 words

There is a particular kind of clarity that comes from sitting across a conference room table at eleven at night, reviewing a client's control documentation with an auditor who is quietly expanding the scope of substantive testing because a reconciliation control failed for the third time in a row. You do not need a textbook to understand, in that moment, what internal controls testing actually means for an organization. It means audit hours, audit fees, and findings that travel up to the audit committee. It means the quality of your controls program is not just a compliance matter; it is the primary variable shaping how much evidence an external auditor decides they need to collect before they will sign their name to your financial statements.

This piece is about that relationship, and about the structured process through which external auditors evaluate it. Understanding the mechanics helps organizations see their controls programs not as bureaucratic infrastructure, but as a direct input to audit scope, findings, and, ultimately, the opinion issued alongside management's assessment to the public.

The Standards and Frameworks That Govern How Auditors Approach Control Testing

In the United States, the primary standard governing integrated audits of public companies is PCAOB Auditing Standard 2201, which requires the auditor to evaluate both the design and the operating effectiveness of internal control over financial reporting alongside the financial statement audit. The two opinions are issued together, dated the same, and neither can be fully understood without the other.

PCAOB AS 2201 tells the auditor what to test and how. The COSO Internal Control – Integrated Framework tells everyone what effective internal control is supposed to look like. COSO's five components, Control Environment, Risk Assessment, Control Activities, Information and Communications, and Monitoring, provide the conceptual skeleton that management uses to design and document its control structure, which the auditor then evaluates against AS 2201's requirements. In practice, you cannot reason about one without the other.

Both are actively evolving. Amendments to AS 2201 are scheduled to take effect December 15, 2026, signaling that the PCAOB views the current standard as needing refinement in response to how auditing practice has changed. Perhaps more immediately consequential is the amended AS 1105, effective for fiscal years beginning on or after December 15, 2025, which requires auditors to explicitly evaluate the reliability of external electronic information used in technology-assisted analysis. As AI-assisted audit techniques proliferate, that requirement elevates the stakes on IT general controls testing across the entire financial reporting supply chain in ways that were not contemplated when the original standard was written.

On the internal audit side, the IIA's 2024 Global Internal Audit Standards, effective January 9, 2025, introduced explicit requirements for documented risk landscape assessments underlying audit plans, and for coordination with external assurance providers. External auditors have long been permitted to rely, with limitations, on work performed by internal audit. These new standards formalize internal audit's side of that relationship, which matters practically because the quality of internal audit's documentation and coordination directly affects how efficiently external auditors can use it.

How Auditors Scope the Engagement and Identify Which Controls to Test

Audit scope is not arbitrary. It follows a top-down risk assessment that begins at the financial statement level: which accounts are significant, which assertions about those accounts carry real risk of material misstatement, and which controls directly address those assertions. The auditor is not trying to test every control in existence; they are trying to identify the controls that, if they failed, would leave a material misstatement undetected.

Entity-level controls function as a powerful scoping lever under AS 2201. If a control operating at the entity level is precise enough to address a specific risk of misstatement on its own, the auditor may not need to test additional process-level controls for that same risk. Organizations with strong, well-documented entity-level controls can meaningfully compress the scope of what gets tested at the transaction or process level for each relevant assertion. That compression has direct consequences for cost and effort on both sides.

The controls that typically land in scope are not exotic: segregation of duties, approvals and authorizations, account reconciliations, access controls, and change management controls over financial systems. These are the workhorses of ICFR, and they are the controls that appear with the greatest frequency in inspection findings precisely because they are tested so often.

AS 2201 permits auditors to use work performed by internal audit and others, but the degree of permissible reliance decreases as the risk associated with a particular control increases. For high-risk controls, the external auditor must perform their own testing. This is a point many organizations misunderstand: internal audit coverage does not substitute for external auditor testing at the places where it matters most.

The practical implication is straightforward. The more rigorously management has documented its risk assessment, its control inventory, and the logic connecting controls to assertions, the more efficiently scoping proceeds. Auditors do not invent scope; they follow the risk. Organizations that have done their own thinking create a more navigable map.

Evaluating Design Effectiveness Before Any Operational Testing Begins

Before asking whether a control worked, an auditor must ask whether it could work. Design effectiveness is the prior question: is this control, if executed as intended, capable of preventing or detecting a material misstatement on a timely basis?

This is not a formality. A control can be executed flawlessly and still fail design evaluation if it is not targeted at the right point in a process, if the person performing it lacks the information or authority to act on what they find, or if its precision is insufficient for the magnitude of risk it is supposed to address. Consistency of execution does not redeem a design that cannot do the job.

Walkthroughs are the primary tool at this stage. An auditor traces a transaction from initiation through authorization, processing, recording, and reporting, observing and asking questions at each step. This is typically done for a single transaction, and it serves a specific and limited purpose: it demonstrates that the auditor understands the process and has confirmed that the control exists and is designed appropriately. A walkthrough establishes design and implementation. It does not, on its own, support any conclusion about whether the control ran effectively across the period.

Implementation is its own element here. The control must exist in practice, not just in a process narrative or a policy document. This distinction matters more than it should, because organizations under time pressure often document controls more meticulously than they execute them, and auditors have learned to look for that gap.

A design deficiency forecloses everything downstream. A control that does not adequately address the risk it is assigned to cannot be rescued by operating effectiveness testing or compensating controls alone. The failure is structural, and it must be remediated at the level of the design before the control can support any reduction in substantive procedures.

Testing Whether Controls Operated Effectively Throughout the Audit Period

Once design is confirmed, the auditor turns to the full audit period. Operating effectiveness asks whether the control functioned as designed, consistently, across every relevant point in time.

The four testing methods are inquiry, observation, inspection of documents, and reperformance. These are not interchangeable. Reperformance, where the auditor independently re-executes the control to verify it produces the correct result, is the strongest, particularly for automated controls where the logic is binary. Inspection of documents, confirming that an approval signature is present, for example, provides evidence but requires the auditor to go further: was the control owner who signed in a position to evaluate the information they were approving, or were they simply executing a formality?

Inquiry alone cannot support a lower control risk assessment. It is corroborative, not sufficient. This is a conceptual distinction that has practical consequences when auditors are under time pressure and inclined to take management's representations at face value.

Automated and manual controls require different testing approaches. An automated control, once its design is confirmed and the surrounding IT general controls are validated, can often be tested once for the period, because the logic does not change between executions. Manual controls require sampling across the period, and sample sizes are driven by risk: higher-risk controls, or controls with exceptions already identified, require larger samples. When exceptions appear in a sample, the auditor's response is not simply to note them; they must assess whether the exception reflects a systemic failure or an isolated deviation, which often requires extending the sample.

The standard requires that if the auditor concludes control risk is anything less than high, they must test operating effectiveness across the period, including roll-forward procedures when interim testing was used. A walkthrough alone is not sufficient to reach that conclusion. This is a point where audit teams occasionally cut corners, and it is precisely the kind of gap PCAOB inspectors are trained to find.

How Control Testing Results Feed Directly Into Substantive Testing Decisions

The link between controls testing and substantive testing is mechanical, not judgmental. Controls that pass testing reduce the auditor's required substantive procedures on the same financial statement assertions; controls that fail require the auditor to compensate with more extensive substantive work to gather sufficient appropriate evidence another way.

Tests of controls assess the reliability of the system. Substantive tests of detail verify individual transactions, balances, and disclosures directly: confirming accounts receivable with customers, vouching invoices to supporting documentation, observing physical inventory counts. Neither type of testing replaces the other entirely, though dual-purpose testing can satisfy both objectives within a single procedure. Auditing standards require some level of substantive procedures regardless of how strong the controls environment appears. But the nature, timing, and extent of those procedures vary significantly based on what the controls testing shows.

That phrase, "nature, timing, and extent," appears across auditing standards and translates directly into audit hours and fees. Changing the extent of substantive testing on a large revenue population, for example, from a targeted sample to a full-population analytical procedure, has measurable consequences for the cost and duration of an audit. Organizations that have experienced an expanded scope following a control failure intuitively understand this; organizations that have not connected the two tend to treat controls investment and audit fees as separate line items.

Revenue and inventory accounts consistently apply pressure at this junction. PCAOB inspection data shows that deficiencies in controls testing cluster repeatedly around these accounts, which then require expanded substantive procedures, which is why those same accounts appear so frequently in inspection findings. The causality runs in both directions: weak controls generate more scrutiny, and more scrutiny generates more findings.

What Deficiency Classification Means and the Consequences That Follow

The deficiency taxonomy has three levels. A control deficiency exists when a control is not designed or operating in a way that would prevent or detect a misstatement. A significant deficiency is more severe, representing a deficiency or combination of deficiencies that is important enough to merit attention from those charged with governance. A material weakness sits at the top: a deficiency, or combination of deficiencies, where there is a reasonable possibility that a material misstatement in the financial statements will not be prevented or detected on a timely basis.

The material weakness standard is not about certainty; it is about reasonable possibility. That threshold is meaningfully lower than most managers expect when they first encounter it, which is why findings sometimes feel disproportionate from the client's perspective.

AS 2201 identifies specific indicators of material weakness. Among them: an auditor's identification of a material misstatement that the company's own internal controls would not have caught before external scrutiny, and ineffective oversight by those charged with governance. The presence of either indicator effectively resolves the severity question.

One or more material weaknesses requires an adverse opinion on ICFR. Not a qualified opinion, not a disclaimer, an adverse one. Because the ICFR and financial statement audits are integrated, that adverse opinion appears alongside the financial statement opinion in the same report. The signal to investors, lenders, and regulators is unambiguous.

Recurrence is more common than most disclosure narratives suggest. Analysis of SEC EDGAR filings through 2025 shows that more than 60% of material weakness disclosures come from repeat filers: organizations that disclosed a weakness, remediated it, and then disclosed again. Remediation that closes the specific control gap without addressing the underlying conditions tends to generate exactly that pattern. A Moss Adams analysis of over 5,000 management assessments across 2020 to 2024 found adverse assessment rates peaked above 26% in 2021, driven substantially by SPAC-related filers, before declining to just over 15% in 2024. The decline is real, but the absolute rate remains high.

Where Auditors Are Actually Finding Gaps: PCAOB Inspection Patterns

The PCAOB's 2024 inspection summary reported a Part I.A deficiency rate of 39%, down from 46% in 2023. Progress is real. Still, more than one in three audits reviewed carried an inspection finding, which is a number worth sitting with.

The improvement is most visible at the largest firms. The Big Four moved from a 26% deficiency rate in both 2022 and 2023 to 20% in 2024. The six largest global network firms combined moved from 34% to 26%. Non-affiliated U.S. firms, by contrast, showed almost no improvement, moving from 53% to 52%. The gap between large and smaller firms is not narrowing in any meaningful way.

The most frequently cited auditing standard in inspection findings, by a substantial margin, is AS 2201. The most frequently cited deficiency is also the most fundamental: auditors did not perform sufficient testing of the design and/or operating effectiveness of controls they themselves selected for testing. This is not a failure of conceptual understanding. Auditors who cannot articulate design effectiveness theory are not passing licensing exams. This is an execution discipline problem, the gap between what the methodology requires and what gets done under time and budget pressure.

The EY 2024 PCAOB inspection (Release No. 104-2025-037) reviewed 64 audits, of which 18 landed in Part I.A. Deficiencies concentrated in revenue and related accounts, inventory, and long-lived assets. The most common failures were insufficient testing of design or operating effectiveness for selected controls, and failure to adequately test controls over the accuracy and completeness of data used in those controls. The KPMG 2024 inspection (Release No. 104-2025-039) reviewed the same number of audits, with 13 in Part I.A, and identified a similar pattern, with the additional citation of auditors failing to identify controls related to a significant account or relevant assertion.

Inventory deserves particular note. Audit deficiencies related to inventory observation appeared across all inspected Big Four firms in 2024, the third consecutive year that failure dominated inspection reports. That kind of consistent, multi-year recurrence across multiple large firms does not point to a technical misunderstanding; it suggests that something structural about how inventory audits are resourced and executed is not working.

The Cost and Operational Reality of Running a SOX Controls Program

According to the KPMG 2025 SOX Survey, the average SOX program now costs $2.3 million annually and requires 15,581 hours, representing a 44% increase in cost and a 32% increase in hours since fiscal year 2022. These are not marginal movements.

The primary driver is scope expansion. The average number of in-scope systems rose 135% over that same period, from 17 to 40. Organizations added technology environments at a pace that outran their ability to design automated controls for them. The counterintuitive result: automated controls actually declined as a share of total controls, from 21% in FY22 to 17% in FY24, even as overall costs rose sharply. More systems, proportionally less automation, and substantially more manual effort.

Hyperproof's 2025 IT and Risk Compliance Benchmark Report found that 59% of respondents now test all controls rather than only critical ones, a 26-percentage-point year-over-year increase. The instinct behind that shift is understandable. After a cycle of inspection findings and material weakness disclosures, broader coverage feels like a safer posture. But broader coverage with less automation and more manual controls does not resolve the execution discipline problem; it amplifies it. More controls tested manually means more opportunity for the kind of inconsistency and documentation failures that generate findings.

The KPMG 2025 analysis of non-IPO companies found material weaknesses rising in 2024 across financial close processes, the control environment, and non-routine or complex transactions. These are areas where human judgment and process discipline are most determinative, and where the combination of expanding scope and declining automation creates the most exposure.

How Technology and AI Are Beginning to Change What Control Testing Looks Like

In 2025, 39% of internal auditors reported using AI tools in their work, according to a survey of over 4,000 audit professionals, with adoption projected to reach 80% by 2026. That rate of change is fast enough to be disorienting even for practitioners who are actively tracking it.

What AI changes, most immediately, is the feasibility of continuous monitoring across full populations. Traditional controls testing relies on sampling, which is efficient but incomplete by design. Technology-assisted analysis allows auditors to flag anomalies across entire transaction datasets that sampling would have a reasonable probability of missing. More practically, it allows reperformance, the strongest testing method, to be applied at a scale that was previously impractical for manual testing.

The amended AS 1105, effective for fiscal years beginning December 15, 2025, requires auditors to explicitly evaluate the reliability of external electronic information used in technology-assisted analysis. As AI-assisted techniques expand, the auditor cannot simply use a data set or a model output without evaluating the controls over the data source itself. This elevates IT general controls testing into a more central role in the audit methodology, not a peripheral one.

COSO issued guidance on generative AI in February 2026, adapting its five components and 17 principles to risks specific to GenAI deployments: hallucination, data leakage, model drift, and unauthorized use. This is a meaningful signal. Control frameworks are beginning to treat AI not only as a tool that auditors use, but as a domain that itself requires controls. Organizations that deploy AI in financial reporting processes, for forecasting, for account analysis, for automated approvals, will need to demonstrate that those systems are controlled with the same rigor as any other element of ICFR.

The underlying logic from every prior section of this piece applies here without modification. Well-governed automated controls and AI systems will support auditor reliance and reduce compensating substantive work. Poorly governed ones will generate the same pattern of expanded scope, more evidence, and harder conversations that poorly governed manual controls have produced. The technology changes the scale of what is possible. It does not change the structure of the problem.

Sources

  1. pcaobus.org
  2. theiia.org
  3. pcaobus.org
  4. hyperproof.io
  5. auditupdate.com
  6. assets.pcaobus.org

More in Financial Statement Audit