How can organisational capability debt be rigorously defined and measured as a…
How can organisational capability debt be rigorously defined and measured as a leading indicator of Artificial Intelligence (AI)-related enterprise risk, and how does pre-existing capability debt amplify risks from autonomous AI systems when human rate limits are removed?
- Capability debt is not a settled literature term, but it can be rigorously operationalised as the accumulated gap between the organisational capabilities required for safe, reviewable, governable AI use and the capabilities actually present in day-to-day practiceC2 (n.d.)Fowler (2009)Institute (2024)National (2023)
- A defensible capability-debt scorecard should combine process appraisal, information-flow culture, governance and inventory coverage, platform and safety-net quality, workaround prevalence, deterministic-control coverage, and workforce skill freshness instead of collapsing risk into a single simplistic metricInstitute (2024)Westrum (2004)National (2023)Google (2025)Macnamara et al. (2024)
- Shadow IT, shadow AI, and citizen-development evidence shows that workaround adoption is usually a demand signal produced by slow or poorly fitted sanctioned capability, which supports treating workaround prevalence as a leading indicator of unmet organisational need and rising governance riskKopper et al. (2020)IBM (2025)McDermott et al. (2025)
- DORA's evidence that AI amplifies existing workflow and platform quality supports treating capability debt as a forward-looking risk signal, because weak safety nets and weak feedback loops become more damaging as AI raises change volume and action frequencyGoogle (2025)
- Pre-existing capability debt becomes a stronger risk amplifier under agentic AI because machine-speed action removes the practical buffering effect of human pace while old management models, review queues, and escalation habits remain too slow to compensateBlog (2026)MIT (2025)Github (n.d.)
- Skill decay should be treated as part of capability debt because AI assistance can weaken human judgment and hide deterioration, which reduces the organisation's ability to review, challenge, and safely contain faster automated output over timeMacnamara et al. (2024)Github (n.d.)Github (n.d.)
- The reviewed frameworks support a sequencing rule in which organisations reduce capability debt first on high-consequence workflows and grant broader autonomy only after explicit controls, inventory, oversight rules, and evaluation evidence are already in placeNational (2023)Google (2025)Blog (2026)MIT (2025)
- Promoting individual AI tools without investing in shared review systems, platform quality, training, and governance creates hidden organisational debt because local productivity rises faster than the collective capacity needed to verify, escalate, and sustain safe useGoogle (2025)IBM (2025)Macnamara et al. (2024)Github (n.d.)
Research Question
How can capability debt, the accumulated organisational deficit in review quality, judgment, process maturity, and skill inventory, be rigorously defined, measured, and tracked as a leading indicator of AI-related enterprise risk? What is the relationship between pre-existing capability debt, including slow central systems, unmet business needs, and weak review culture, and the amplification of those risks when autonomous, goal-directed AI systems (agentic AI) remove human rate limits? How should organisations sequence debt reduction relative to AI rollout, and in what ways does promoting individual AI tools without corresponding investment in review and quality systems create hidden organisational debt?
Findings
Executive Summary
The accumulated shortfall between the organisational capabilities required for safe AI scale and the capabilities actually present in practice should be tracked as a leading indicator of future AI-related risk rather than as a lagging description of incidents that have already happened.
Capability debt is used here as shorthand for that shortfall, because the reviewed sources provide the component parts of the construct but do not standardise the label itself.
Agentic AI, used here to mean autonomous, goal-directed AI systems that can plan and execute actions with limited continuous human oversight, amplifies pre-existing capability debt because the same organisations that already rely on workarounds or weak review culture are then asked to govern machine-speed action with unchanged or deteriorating review capacity.
The strongest practical conclusion is a sequencing rule, reduce capability debt first on high-consequence control surfaces and allow broader autonomy only where deterministic controls, inventory, oversight rules, and evaluation evidence already exist.
Key Findings
- Capability debt is not a settled literature term, but it can be rigorously operationalised as the accumulated gap between the organisational capabilities required for safe, reviewable, governable AI use and the capabilities actually present in day-to-day practice.
- A defensible capability-debt scorecard should combine process appraisal, information-flow culture, governance and inventory coverage, platform and safety-net quality, workaround prevalence, deterministic-control coverage, and workforce skill freshness instead of collapsing risk into a single simplistic metric.
- Shadow IT, shadow AI, and citizen-development evidence shows that workaround adoption is usually a demand signal produced by slow or poorly fitted sanctioned capability, which supports treating workaround prevalence as a leading indicator of unmet organisational need and rising governance risk.
- DORA's evidence that AI amplifies existing workflow and platform quality supports treating capability debt as a forward-looking risk signal, because weak safety nets and weak feedback loops become more damaging as AI raises change volume and action frequency.
- Pre-existing capability debt becomes a stronger risk amplifier under agentic AI because machine-speed action removes the practical buffering effect of human pace while old management models, review queues, and escalation habits remain too slow to compensate.
- Skill decay should be treated as part of capability debt because AI assistance can weaken human judgment and hide deterioration, which reduces the organisation's ability to review, challenge, and safely contain faster automated output over time.
- The reviewed frameworks support a sequencing rule in which organisations reduce capability debt first on high-consequence workflows and grant broader autonomy only after explicit controls, inventory, oversight rules, and evaluation evidence are already in place.
- Promoting individual AI tools without investing in shared review systems, platform quality, training, and governance creates hidden organisational debt because local productivity rises faster than the collective capacity needed to verify, escalate, and sustain safe use.
Assumptions
- A composite scorecard is more decision-useful than a single score because different debt classes fail in different ways and need different interventions.
- Westrum's healthcare safety-culture typology generalises sufficiently to enterprise AI governance because both settings depend on escalation quality, information flow, and the treatment of bad news.
- Some low-risk debt can be tolerated in bounded advisory use cases because the reviewed frameworks support progressive, risk-tiered autonomy rather than all-or-nothing deployment decisions.
Analysis
The evidence weighs most strongly in favour of treating capability debt as an operational synthesis construct rather than as an already-standardised academic term.
That synthesis is still rigorous because each component of the construct is independently evidenced: process maturity is appraisable, information culture predicts safety performance, AI governance requires inventory and oversight, workflow quality determines whether AI amplification is stabilising or destabilising, and workaround prevalence reveals unmet demand.
The competing interpretation is that organisations should simply deploy AI quickly and rely on later governance hardening, but the reviewed sources point the other way on high-consequence surfaces because they repeatedly require explicit controls, lifecycle oversight, and safety nets before broad autonomy is expanded.
The strongest rival remedy is to preserve traditional human review rather than reducing capability debt, but MIT Sloan and AWS both warn that generic human approval collapses when volume rises, which means staffing alone does not solve the structural gap unless review rules, thresholds, skills, and external controls are also redesigned.
Risks, Gaps, and Uncertainties
- No reviewed source provides a validated off-the-shelf capability-debt index tied directly to later AI incident rates.
- Shadow-AI and LCNC evidence is strong on drivers and prevalence, but much of it remains survey-based or synthesis-based rather than longitudinal causal measurement.
- The skill-decay source is theoretically strong but does not yet provide enterprise-scale incident correlations for AI reviewer populations.
- Agentic-AI governance literature is moving quickly, so some sequencing guidance will likely become more explicit over the next review cycle.
Open Questions
- Which scorecard thresholds best separate tolerable from intolerable capability debt for specific workflow classes such as advisory, read-only, and write-capable operations?
- Which workaround signals most reliably distinguish healthy local experimentation from evidence of systemic sanctioned-path failure?
- What reviewer-practice regime preserves human judgment best once agents are operating continuously at enterprise scale?
sources
- [x] Ward Cunningham (1992) The WyCash Portfolio Management System - original debt-metaphor text used to anchor what debt means before extending the metaphor to organisational capability.
- [x] Martin Fowler (2009) Technical Debt Quadrant - deliberate versus inadvertent, prudent versus reckless debt taxonomy used to structure the analogy.
- [x] ISACA CMMI Institute (2024) What is an Appraisal? - accessible process-appraisal source for benchmarked capability measurement.
- [x] Westrum (2004) A typology of organisational cultures - information-flow and safety-culture typology used as a leading-indicator lens.
- [x] National Institute of Standards and Technology (NIST) (2023) AI Risk Management Framework Core - governance, inventory, training, human oversight, and monitoring outcomes relevant to capability measurement.
- [x] Google Cloud DevOps Research and Assessment (DORA) (2025) Announcing the 2025 DORA report - AI-as-amplifier and platform-quality prerequisite evidence.
- [x] Amazon Web Services (AWS) Security Blog (2026) Four security principles for agentic AI systems - machine-speed autonomy, deterministic external controls, and earned-autonomy sequencing.
- [x] MIT Sloan Management Review and Boston Consulting Group (BCG) (2025) Agentic AI at Scale: Redefining Management for a Superhuman Workforce - explicit claim that old human-paced management models break under superhuman speed and scale.
- [x] IBM (2025) Is rising AI adoption creating shadow AI risks? - employee demand, external-tool substitution, and training or data-infrastructure gaps behind shadow AI.
- [x] Kopper et al. (2020) From Shadow Information Technology (IT) to Business-managed IT - empirical shadow-Information Technology (IT) evidence that workarounds appear when central delivery cannot provide suitable systems quickly enough.
- [x] McDermott et al. (2025) Adoption of low-code and no-code development: a systematic literature review and future research agenda - systematic review of citizen-development and Low-Code/No-Code (LCNC) adoption.
- [x] Macnamara et al. (2024) Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers' awareness? - evidence that AI assistance can erode judgment and obscure deterioration in human skill.