Est.

Internal Controls Testing During an Audit

Design flaws and execution gaps require different testing methods to catch reliably.

Senior Writer · · 17 min read
Cover illustration for “Internal Controls Testing During an Audit”
Financial Statement Audit · July 29, 2026 · 17 min read · 3,716 words

Design evaluation answers a threshold question: even if this control runs perfectly, would it catch the risk it is supposed to catch? If the answer is no, nothing that follows matters.

Auditors evaluate design through three primary means. Inquiry involves direct conversations with control owners. Observation means watching the control execute in real time. The walkthrough traces a transaction from initiation through recording, confirming that controls are present and correctly positioned at each handoff. Of the three, the walkthrough carries the most diagnostic weight, though the way most practitioners conduct them, you would not necessarily know it.

The walkthrough gets treated as documentation theater far more often than it should. What it actually surfaces is the gap between what a policy document says and how work moves through a process, and that gap is frequently wider than anyone expects. A written procedure may describe a three-way match before payment approval; the walkthrough reveals that the matching step is skipped for vendors flagged as trusted suppliers, with no compensating control in place. That gap is invisible from the policy alone. I have found this kind of divergence at organizations with well-resourced, professionally staffed finance functions. That is exactly what makes it easy to miss if you walk in assuming competence implies compliance. It usually does, until it doesn't.

Design failure modes follow recognizable patterns once you have seen enough of them: controls set at a level too aggregated to catch individual transaction errors; segregation-of-duty gaps built into the process by design rather than treated as exceptions; approval thresholds set above the materiality level for the risk they nominally address; preventive controls substituted for detective controls in areas where timing of detection is itself the exposure. A control requiring manager approval for purchases over fifty thousand dollars does not address misappropriation occurring through dozens of transactions below that threshold. The design, however consistently executed, simply does not cover the exposure. Consistent operation of a broken control is not a mitigant.

What passes design evaluation advances to operating effectiveness testing. What fails gets flagged as a design deficiency, and the auditor plans around it by expanding substantive work in the affected area. There is no remediation path within the same audit period.

Testing whether controls work in practice, not just on paper

Operating effectiveness testing asks whether the control ran as designed, by the right people, at the right frequency, throughout the period under audit. The five primary methods available carry different evidentiary weight and impose different constraints on what can actually be concluded.

Inquiry is necessary but sufficient on its own only in combination with other methods. Management confirming that a control works is not evidence that it does; it is a starting point for corroboration. Observation is direct, but it captures a single instance, which limits what you can say about a full year of performance. Inspection covers documentation of past performance: approval signatures, system logs, reconciliation records. Re-performance means the auditor independently repeats the control procedure and compares results to the client's output, and it carries the highest evidentiary weight, particularly for calculation-based and automated procedures. Computer-assisted techniques allow auditors to work across large populations in ways that were not feasible even ten years ago.

The strength of any operating effectiveness conclusion comes from combining methods so they corroborate each other. Inquiry without inspection is anecdote. Observation without re-performance confirms presence but not accuracy. The nature of the control shapes which combination is appropriate: manual controls require observation and documentation inspection; IT-dependent controls require re-performance and system data review; automated controls can be tested with limited instances once the underlying logic is verified, though that verification depends entirely on the integrity of the IT environment.

Frequency shapes what a sufficient test looks like in a way that practitioners sometimes underestimate. A monthly reconciliation runs twelve times a year; a daily transaction approval may run hundreds of times. The higher the frequency, the larger the population that needs to be considered, and the more instances are required to support a defensible conclusion about consistency. That relationship sounds obvious, and it is routinely underweighted in practice.

How sampling decisions shape what the test can actually prove

Sampling in controls testing is neither guesswork nor a rigid formula. Sample size reflects a combination of factors: control frequency, assessed risk, tolerable deviation rate, and whether the auditor is relying on results to reduce substantive work. The general logic scales with frequency, and the tolerable deviation rate set during planning anchors how much that logic can bend. Annual controls typically require one or two instances because the entire population is that small. Monthly controls call for a handful. Daily or transaction-level controls require significantly larger samples, often twenty-five or more, because even a modest deviation rate across thousands of instances represents a meaningful failure pattern.

A clean sample supports a conclusion of operating effectiveness. It does not guarantee that no failures occurred during the period. Sampling provides reasonable assurance within defined parameters, not certainty, and the difference between those two things matters when documenting reliance decisions or explaining conclusions to a client who wants more than you can responsibly give.

Population definition deserves more attention than it typically receives. The failure mode is subtle, which is part of why it persists. The population is defined too narrowly, perhaps restricted to high-dollar transactions when the control is supposed to cover all transactions. The sample comes back clean. The conclusion is technically supported by what was tested. And yet the test does not address the risk in question. The working papers look fine until someone asks the right question. I have watched careful auditors, people whose rigor I respected in other respects, produce conclusions that could not survive that question, precisely because the population definition was made early in planning and never revisited as their understanding of the control evolved. The assumption baked in at the start became load-bearing without anyone noticing.

Selection method carries its own documentation burden. True random sampling, the basis of attribute sampling under statistical sampling standards, provides statistical defensibility. Judgmental selection is permissible under professional standards but requires explicit documentation and cannot be used to select only favorable instances. Haphazard selection that systematically avoids known problem periods is not a methodology; it is a bias with working papers attached.

When deviations appear, the auditor faces a branching decision: expand the sample to determine whether the deviation is isolated or part of a pattern; reassess the level of reliance that can be placed on the control; or shift to a primarily substantive approach. Each path has downstream implications for scope, timing, and resource requirements, and each requires documented reasoning, not just a conclusion.

The difference between a deficiency, a significant deficiency, and a material weakness

Control failures are not equivalent, and the classification of a failure determines what must be communicated, to whom, and in what form. Underclassifying a material weakness protects no one. Overclassifying a minor gap creates unnecessary alarm and erodes the credibility that governance communication depends on.

A control deficiency exists when a control is absent, poorly designed, or ineffectively operated, but the shortfall does not rise to a level that materially threatens the financial statements. These are documented in working papers and may warrant a management letter, but they do not require escalation to governance. A significant deficiency is important enough to warrant the attention of those charged with governance even if it does not create a reasonable possibility of material misstatement; the communication obligation is formal and written. A material weakness exists when a deficiency, or combination of deficiencies, creates a reasonable possibility that a material misstatement will not be prevented, detected, or corrected on a timely basis. For public companies, this finding lands in the auditor's report on internal control over financial reporting.

The aggregation of deficiencies is where classification judgment gets particularly hard. Individual deficiencies that appear modest in isolation can combine into a significant deficiency or material weakness when they cluster around the same risk area or affect the same financial statement line. A gap in access controls, a weakness in the reconciliation process, an override pattern in the approval workflow: each may be individually tolerable. Together, covering the same revenue cycle, they may constitute a material weakness. One might argue that evaluating each deficiency on its own merits is sufficient, but stopping there is itself a classification error, and not a conservative one.

Aggregation analysis should not be a wrap-up exercise. It needs to run as a continuous thread throughout fieldwork, updated as each new finding comes in, because a late-breaking addition to the deficiency list may change the classification of everything that preceded it. I have seen that happen. The instinct is to treat the new finding on its own terms; the correct instinct is to go back and re-examine what you already concluded.

PCAOB AS 2201, amended with an effective date of December 15, 2026, governs the evaluation and reporting of internal control deficiencies for public company integrated audits. The standard requires classification judgments to be documented with explicit reasoning. Working papers that cannot reconstruct the logic of a classification decision do not meet current requirements, let alone the amended version. Firms auditing issuers would be better served preparing now than retrofitting their workflows under deadline pressure.

How controls testing results determine the scope of substantive work

The audit risk model is the conceptual architecture connecting controls testing to everything else. Audit risk is a function of inherent risk, control risk, and detection risk. Detection risk is the only variable the auditor actively manages, reflecting the probability that audit procedures will fail to detect a material misstatement that exists. When controls testing supports reliance, the auditor can accept higher detection risk in the affected area and do proportionally less direct testing. When it does not, the auditor shifts to a primarily substantive approach: more transactions examined directly, larger samples, more extensive analytical procedures.

Dual-purpose testing is an underused efficiency lever here. Some procedures simultaneously serve as both a test of controls and a substantive test. Examining a sample of invoices can confirm that an approval control operated and provide direct evidence about the recorded amounts. Planning dual-purpose tests from the outset, rather than treating controls and substantive work as entirely separate streams, compresses timelines without sacrificing rigor. The conversation about where dual-purpose testing is feasible belongs in planning, not fieldwork. By fieldwork, the structure of the work is largely set.

Interim versus year-end testing introduces a timing dimension with real exposure. Auditors frequently test controls at an interim date, often three to six months before year-end, and carry reliance forward. That is acceptable under professional standards only if the auditor has assessed whether the control environment changed materially between the interim date and period end. A significant personnel departure, a system migration, or a change in the approval process during the roll-forward period can invalidate interim reliance entirely, and roll-forward procedures that are planned but not executed represent one of the more avoidable exposure points in the interim-to-year-end transition. The errors I have seen here are mundane rather than dramatic: roll-forward procedures planned but not executed; changes identified but insufficiently evaluated. Both are avoidable and both more common than they should be.

There are areas where even strong controls rarely reduce substantive work to a minimum. Accounting estimates, fair value measurements, loss reserves, and related-party transactions require substantive procedures regardless of control strength, because the risk lives not in the mechanical execution of a control but in the judgment underlying the number itself. Controls cannot fully govern judgment, and professional standards reflect that reality consistently.

Where IT controls and automated controls fit into the testing framework

Automated controls introduce an asymmetry relative to manual controls that looks like a simplification until you follow it to its dependencies. Once an automated control has been verified as correctly programmed and operating within an unchanged IT environment, it can generally be tested with a small number of instances. The logic runs consistently by design; if it worked correctly on one transaction, it worked correctly on all transactions processed under the same conditions.

The word "unchanged" carries most of the weight in that sentence. The entire basis for limited-instance testing of automated controls is the assumption that the IT environment surrounding them is stable and well-governed, and that assumption holds only when IT general controls are operating effectively. ITGCs cover access controls (who can read, write, or approve within a system), change management (how system changes are authorized, tested, and deployed), IT operations (job scheduling, error handling, data backups), and program development controls governing how new applications and interfaces enter the production environment. These are not supplementary concerns; they are prerequisites.

A material weakness in ITGCs can cascade across every automated control that depends on the same systems. A single failure point in access controls over a core financial system multiplies audit risk across the entire population of transactions processed by that system. Auditors who test automated controls without first establishing the integrity of the underlying ITGC environment are constructing conclusions on an untested foundation. Organizations that have expanded their reliance on automated controls in recent years without a corresponding investment in ITGC coverage have quietly accumulated risk that the efficiency narrative of automation tends to obscure. That tension does not resolve itself.

Computer-assisted audit techniques and data analytics now allow auditors to test entire populations for certain control attributes rather than relying on samples. They surface exceptions that sampling would likely miss: transactions split below approval thresholds, anomalous access patterns, approval workflow bypasses. The PCAOB's amendments to AS 1105, effective for fiscal years beginning on or after December 15, 2025, require explicit evaluation of the reliability of external electronic information used in technology-assisted analysis. Auditors using data analytics need to demonstrate that the populations they ingested were complete, accurate, and sourced from a reliable system. That demonstration cannot be assumed; it has to be documented.

Continuous controls monitoring represents a further evolution. Rather than performing controls testing as a point-in-time exercise during fieldwork, continuous monitoring integrates analytics into the IT environment to flag deviations as they occur, shifting the assurance model from retrospective sampling to ongoing detection. The operational and documentation implications of that shift are still being worked out across the profession.

How AI is beginning to change what controls testing looks like in practice

The efficiency gains AI is creating in controls testing are real, but they are concentrated in specific tasks, and understanding where those tasks end matters as much as understanding where they begin.

The clearest applications are where testing has historically been constrained by data volume: exception testing across full transaction populations, automated matching of supporting documentation to recorded amounts, anomaly detection in access and approval patterns. Procurement data can be analyzed across entire populations for transactions split below approval thresholds. User access logs can be scanned for patterns that would take a human auditor hours to surface manually. Agentic AI tools can execute defined testing procedures, including evidence extraction, population filtering, and exception flagging, within parameters set by the practitioner.

The productivity signal from firms actively deploying these tools is notable. Rightworks' 2025 Accounting Firm Technology Survey found that firms using AI reported thirty-seven percent higher revenue per employee compared to non-adopters, with time savings and task automation as the top reported benefits. That gap reflects what becomes possible when technology coverage expands faster than headcount.

What AI does not do is transfer the auditor's judgment about reliability and sufficiency to the tool. An AI tool that flags a set of exceptions has surfaced a population for review. The auditor still has to determine whether the exceptions are actual deviations, understand why they occurred, and conclude on what they mean for reliance. The analysis remains a professional responsibility; the tool accelerates the legwork. Whether sufficiently advanced AI could eventually absorb some of that judgment is a question the profession cannot yet answer cleanly. Model explainability, auditability of the tool itself, and regulatory acceptance remain open problems. Confident predictions in either direction are difficult to justify.

The validation discipline this requires is specific. Testing AI tools against completed prior-year audits before deploying them on live engagements provides a basis for evaluating accuracy. Piloting on live engagements with an explicit, documented learning period builds the defensibility that inspection-ready work requires. Data quality is the binding constraint throughout: AI tools processing poorly structured or incomplete client populations produce unreliable outputs, not faster ones. The auditor needs to understand the data before the tool touches it. Treating these tools as set-and-forget solutions creates compliance exposure that tends to surface at the worst possible time, which is to say, during the engagement itself.

What the current regulatory environment means for how controls testing is documented and reported

Documentation is where execution becomes defensible, or fails to. The standard for controls testing working papers is consistent across frameworks: what was tested, by what method, over what population, with what result, and what conclusion was drawn. A conclusion that cannot be reconstructed from the workpaper is not a conclusion; it is an assertion. That is the most common failure mode I have encountered across peer reviews and inspections, more common than any technical testing error, and it is almost a time-pressure problem rather than a knowledge problem.

PCAOB AS 2201, amended with an effective date of December 15, 2026, updates requirements for evaluating and communicating internal control deficiencies in integrated audits of public companies. Classification judgments must be documented with explicit reasoning. Auditors who cannot reconstruct that reasoning from their working papers are already falling short of current requirements; the amended standard simply makes the consequences more visible.

QC 1000, also effective December 15, 2026, requires firms to maintain risk-based quality control systems that proactively identify and manage risks to audit quality, mirroring the structure of ISQM 1 for firms with international engagements. The structural shift is from post-hoc inspection to built-in oversight that surfaces risks before they become findings. For controls testing, that means quality checkpoints woven into the workflow rather than reviewed only at sign-off. Whether firms are substantively restructuring their quality processes or updating documentation to reflect the appearance of doing so will become clearer as the effective date approaches and inspection results accumulate. The distinction matters more than the standard itself will be able to enforce.

The New IIA Global Internal Audit Standards, effective January 9, 2025, introduced mandatory Topical Requirements for AI, cybersecurity, and third-party risk. Audit programs providing assurance over these areas must explicitly reflect those requirements. A prior-year program that was adequate may fall short now, and the determination of adequacy has to be documented rather than assumed.

COSO's February 2026 guidance on internal control over generative AI adapted the five COSO components to GenAI-specific risks, including hallucination, data leakage, model drift, and unauthorized use. Organizations that have deployed generative AI in finance functions have created exposures that existing controls frameworks were not designed to address. For audit teams whose clients have already deployed these tools, that is a present concern, not a future one.

The SEC enforcement environment has shifted in tone under current leadership, with signals that books-and-records or internal controls violations will not automatically be treated on par with fraud. That is a prioritization signal. Reading enforcement posture as a basis for relaxed documentation expectations would be a misreading with significant downside risk.

What a well-structured controls testing program looks like end to end

A controls testing program that holds up under scrutiny is not built during fieldwork. The quality of the planning determines whether fieldwork produces defensible conclusions or documented activity, and additional testing after the fact does not fully compensate for a poorly structured plan.

Planning begins with a risk assessment that maps financial statement assertions to specific control objectives: which controls are relevant to which assertions, over which accounts, at which organizational levels. This mapping prevents the two most common structural failures, testing controls that are irrelevant to the assessed risks and failing to test controls that are relevant to the assertions actually at risk in each account balance. The universe of relevant controls should be identified and agreed upon before a single sample is pulled.

From that mapping, the auditor determines which controls merit reliance testing. Not every relevant control needs to be tested for operating effectiveness if the strategy is primarily substantive in a given area. The decision about where to invest in reliance testing versus where to proceed directly to substantive work is a planning judgment and should be documented with the reasoning behind it.

For controls selected for testing, the plan specifies the method, the population definition, the sample size logic, and the criteria for evaluating results, all before testing begins. Defining what constitutes a deviation, and what deviation rate would undermine reliance, in advance of seeing the results is not bureaucracy. It is what prevents the conclusion from being shaped by what the sample happened to produce. That discipline is difficult to maintain under schedule pressure, and it also matters most precisely then.

Execution follows the plan, with deviations from the plan documented as they occur. When exceptions appear, they are evaluated against pre-defined criteria. Expansion decisions, reliance reassessments, and fallbacks to substantive procedures are recorded with the reasoning behind them, not just the outcome.

Deficiency evaluation runs as a continuous thread. As testing produces results, findings are aggregated across control areas to assess whether combinations rise to the level of significant deficiency or material weakness. Waiting until fieldwork is complete to evaluate aggregation creates timeline risk and reduces the analytical quality of the assessment.

Communication closes the loop. Management receives timely communication of deficiencies as they are identified. Significant deficiencies go to the audit committee in writing. Material weaknesses appear in the auditor's report. Each communication is supported by workpapers that document the basis for the classification.

What makes a program well-structured is not its complexity. Its defining quality is traceability: the planning logic flows through to the testing approach, the testing results flow through to the conclusions, and the conclusions are documented in a way that can be evaluated by someone who was not present. Without that, a controls testing program does not actually manage audit risk. It creates a record of having engaged with it, which is a different thing entirely.

Sources

  1. zengrc.com
  2. diligent.com
  3. ecampusontario.pressbooks.pub
  4. fieldguide.io
  5. pcaobus.org
  6. onspring.com

More in Financial Statement Audit