AI for Control Testing, Gap Identification, and Policies/Standards Reviews

2026-03-03 · agentic-ai rag-retrieval governance-policy security-risk regulatory-compliance · medium · source → · wiki →
key claims
  1. GRC platforms deliver continuous control monitoring in production. AuditBoard's Accelerate (launched 2025) automates sample selection, evidence gathering, and workpaper generation for internal audit teams. ServiceNow GRC deploys AI agents for end-to-end compliance workflows, continuous monitoring, and anomaly detection. Vanta and Drata automate up to 90% of evidence collection with hourly test execution across SOC 2, ISO 27001, HIPAA, and PCI-DSS frameworks
  2. All four Big-4 firms have embedded AI in their external audit platforms. KPMG Clara integrates generative AI for risk assessment, substantive testing, and documentation review, paired with MindBridge transaction scoring for anomaly detection. EY Helix deploys GL and cycle analyzers for full-population transactional testing. PwC Halo for Journals uses ML to flag unusual transactions. Deloitte Omnia deploys GenAI and agentic AI for documentation review, financial statement navigation, and memo drafting. All four maintain human-in-the-loop requirements for final audit conclusions
  3. Regulatory gap analysis has a mature specialist vendor market. Deloitte's automated gap analysis tool uses GenAI to compare regulatory text (DORA, EU AI Act, CRR) against internal policies, outputting structured gap reports. Kodex AI maps DORA, PSD3, and custom frameworks against internal policies with regulatory traceability. Trustero completes full gap assessments in under two hours against DORA and NIST, replacing consulting projects that previously took months. RiskCognition reports >92% accuracy versus manual review in European bank DORA deployments
  4. IAASB formally adopted a Technology Position in October 2024. The position mandates ongoing monitoring and updating of standards to address AI and machine learning in audit. Under existing ISAs, AI-generated evidence is acceptable if it meets the "sufficient and appropriate" standard (ISA 500). Auditors must validate tools before use, apply professional skepticism to outputs, and document their assessment proportionally to tool complexity. ISQM 1 quality management standards govern firm-level certification of automated tools
  5. PCAOB is updating standards to address AI in audit. In 2024 the PCAOB issued a spotlight publication on AI use and signalled amendments to AS 1105 (audit evidence) and AS 2301 (audit procedures). The aggregate Part I.A deficiency rate for Big 4 firms fell to 20% in 2024 (from 26% in 2023), partially attributable to technology adoption. The PCAOB's position is that AI supplements but does not replace human auditors; professional skepticism and quality controls must encompass AI tool outputs
  6. No standard-setter has prohibited AI-generated assurance evidence; all require human oversight. The FRC published landmark guidance in 2025 providing illustrative examples of AI-enabled audit evidence and required certification/testing processes. The universal governance model is: AI executes testing and generates outputs; a qualified human reviews, applies judgment, and takes ownership of the conclusion. Delegation of the conclusion itself to AI — with no human review — is not accepted under any current framework
  7. IIA's September 2024 AI Auditing Framework provides the practitioner governance model. The updated framework covers AI strategy, cyber risks, vendor controls, ethics, bias, and staff training, organised along the Three Lines Model. It addresses both organisations auditing AI and organisations deploying AI in internal audit itself. AI use in internal audit more than doubled in adoption (15% to 40% within a year per IIA data), with upskilling identified as the primary constraint on further adoption
  8. Disclosed practitioner case studies show meaningful efficiency gains. Baker Tilly documented a financial institution using GenAI to automate compliance testing, extracting unstructured data and mapping it into compliance models with real-time monitoring. RSM documented a global bank using intelligent automation for control testing, shifting auditors from data gathering to risk-centric fieldwork. Global banks have implemented centralised AI model inventory platforms, enabling transparency and auditability of AI decision-making across compliance, risk, and audit functions

Research Question

Which organisations are using AI to automate control testing, identify control gaps, or conduct policies and standards reviews — and what does the current vendor, practitioner, and regulatory landscape look like for AI-assisted assurance in financial services?

Findings

Executive Summary

AI-assisted control testing and regulatory gap analysis have moved from aspiration to active commercial deployment. Every major GRC platform (AuditBoard, ServiceNow, Diligent, LogicGate) and all four Big-4 audit firms (KPMG Clara, EY Helix, PwC Halo/Aura, Deloitte Omnia) have production AI capabilities for automated control testing, workpaper generation, and continuous monitoring as of 2024–2025. The regulatory standard-setter position — from IAASB, PCAOB, and FRC — is that AI-generated audit evidence is acceptable provided it is validated, documented, and subject to human professional judgment; no regulator has prohibited it. RBNZ has flagged systemic risks from AI adoption broadly but has not issued prescriptive guidance on AI-assisted assurance; the intersection with BS11 outsourcing obligations remains the primary governance constraint for NZ-supervised entities.

Key Findings

  1. GRC platforms deliver continuous control monitoring in production. AuditBoard's Accelerate (launched 2025) automates sample selection, evidence gathering, and workpaper generation for internal audit teams. ServiceNow GRC deploys AI agents for end-to-end compliance workflows, continuous monitoring, and anomaly detection. Vanta and Drata automate up to 90% of evidence collection with hourly test execution across SOC 2, ISO 27001, HIPAA, and PCI-DSS frameworks.

  2. All four Big-4 firms have embedded AI in their external audit platforms. KPMG Clara integrates generative AI for risk assessment, substantive testing, and documentation review, paired with MindBridge transaction scoring for anomaly detection. EY Helix deploys GL and cycle analyzers for full-population transactional testing. PwC Halo for Journals uses ML to flag unusual transactions. Deloitte Omnia deploys GenAI and agentic AI for documentation review, financial statement navigation, and memo drafting. All four maintain human-in-the-loop requirements for final audit conclusions.

  3. Regulatory gap analysis has a mature specialist vendor market. Deloitte's automated gap analysis tool uses GenAI to compare regulatory text (DORA, EU AI Act, CRR) against internal policies, outputting structured gap reports. Kodex AI maps DORA, PSD3, and custom frameworks against internal policies with regulatory traceability. Trustero completes full gap assessments in under two hours against DORA and NIST, replacing consulting projects that previously took months. RiskCognition reports >92% accuracy versus manual review in European bank DORA deployments.

  4. IAASB formally adopted a Technology Position in October 2024. The position mandates ongoing monitoring and updating of standards to address AI and machine learning in audit. Under existing ISAs, AI-generated evidence is acceptable if it meets the "sufficient and appropriate" standard (ISA 500). Auditors must validate tools before use, apply professional skepticism to outputs, and document their assessment proportionally to tool complexity. ISQM 1 quality management standards govern firm-level certification of automated tools.

  5. PCAOB is updating standards to address AI in audit. In 2024 the PCAOB issued a spotlight publication on AI use and signalled amendments to AS 1105 (audit evidence) and AS 2301 (audit procedures). The aggregate Part I.A deficiency rate for Big 4 firms fell to 20% in 2024 (from 26% in 2023), partially attributable to technology adoption. The PCAOB's position is that AI supplements but does not replace human auditors; professional skepticism and quality controls must encompass AI tool outputs.

  6. No standard-setter has prohibited AI-generated assurance evidence; all require human oversight. The FRC published landmark guidance in 2025 providing illustrative examples of AI-enabled audit evidence and required certification/testing processes. The universal governance model is: AI executes testing and generates outputs; a qualified human reviews, applies judgment, and takes ownership of the conclusion. Delegation of the conclusion itself to AI — with no human review — is not accepted under any current framework.

  7. IIA's September 2024 AI Auditing Framework provides the practitioner governance model. The updated framework covers AI strategy, cyber risks, vendor controls, ethics, bias, and staff training, organised along the Three Lines Model. It addresses both organisations auditing AI and organisations deploying AI in internal audit itself. AI use in internal audit more than doubled in adoption (15% to 40% within a year per IIA data), with upskilling identified as the primary constraint on further adoption.

  8. Disclosed practitioner case studies show meaningful efficiency gains. Baker Tilly documented a financial institution using GenAI to automate compliance testing, extracting unstructured data and mapping it into compliance models with real-time monitoring. RSM documented a global bank using intelligent automation for control testing, shifting auditors from data gathering to risk-centric fieldwork. Global banks have implemented centralised AI model inventory platforms, enabling transparency and auditability of AI decision-making across compliance, risk, and audit functions.

  9. RBNZ has flagged AI systemic risks but has not issued prescriptive assurance guidance. The May 2025 Financial Stability Report "Rise of the Machines" identifies third-party AI concentration risk, model transparency, and data governance as primary concerns. RBNZ expects regulated entities to maintain transparency and auditability of AI outputs, validate results before they inform compliance actions, and manage AI vendor dependency under BS11 outsourcing obligations. No specific standard for AI-generated audit evidence has been issued.

  10. The governance gap for Type 3 agents (delegated authority) remains open. Current frameworks are explicit that AI can generate testing workpapers, sample recommendations, and gap analyses (Type 2), but the final determination that a control is "operating effectively" must be made or owned by a qualified human. The governance question — what reviewer competency is required, what documentation suffices, and whether the determination constitutes an outsourced function under BS11 — is not yet resolved by any regulator.

  11. NZ-supervised entities can access the full vendor and Big-4 landscape. All four Big-4 firms operate in NZ and their global AI audit platforms are deployed here. ServiceNow, AuditBoard, LogicGate, Vanta, and Drata are all available to NZ enterprises. Specialist DORA gap analysis tools have less immediate NZ relevance (DORA is EU-specific), but the equivalent analysis against RBNZ standards (BS11, BS2A) could be built on the same tooling approach.

  12. Full-population testing is displacing sample-based testing in automated control environments. Where AI can access complete transaction populations, audit approaches have shifted from statistical sampling to exhaustive testing. This improves control assurance but raises new questions about what constitutes a meaningful control test and what the auditor adds when the system executes and evaluates every transaction.

Assumptions

Analysis

The vendor and practitioner landscape has bifurcated. Horizontal GRC platforms (AuditBoard, ServiceNow, Vanta, Drata, LogicGate) focus on continuous automated control monitoring — evidence collection, exception flagging, and test execution at scale. Specialist gap analysis tools (Deloitte, Kodex AI, Trustero, RiskCognition) focus on comparing regulatory text against internal policy inventories using NLP. Big-4 audit firms sit across both: they use proprietary platforms for external audit (Clara, Helix, Halo, Omnia) and have built or licensed specialist gap analysis tools for advisory engagements.

The regulatory acceptability question has been answered in principle but not in detail. Every major standard-setter has confirmed that AI-generated evidence is acceptable under existing frameworks, provided it is validated, documented, and subject to human professional judgment. The outstanding gap is granularity: what competency must the reviewing human have? What documentation satisfies professional skepticism requirements when the AI generated the test procedure, executed it, and drafted the conclusion? These questions are being addressed incrementally through guidance (FRC 2025, IAASB 2024–2026 work programme) rather than standard revision.

The RBNZ position reflects the financial stability angle rather than an assurance standard position. RBNZ is concerned about systemic risk from AI concentration and about the reliability of AI model outputs affecting financial decisions — not specifically about whether AI-generated audit workpapers meet evidence standards. For NZ-supervised entities, the practical constraint is BS11 outsourcing governance: if AI control testing is delivered as a third-party service (SaaS), the outsourcing risk management programme must cover it. This is a well-understood governance requirement, not a novel barrier.

The most consequential open question is governance for Type 3 agents: AI that makes the determination that a control is operating effectively, not just AI that generates the evidence for a human to assess. Current frameworks stop short of endorsing this. The practitioner trajectory — moving from sampling to full-population testing, from workpaper generation to conclusion drafting — is pushing toward Type 3 faster than governance frameworks are adapting. The next 12–24 months will likely see standard-setters (IAASB, PCAOB) issue more specific guidance on what human review must consist of when AI has performed the entire test cycle.

Risks, Gaps, and Uncertainties

Open Questions


sources


Connected items

Loading…

View full knowledge graph →