Q6: Leading indicators of instability in split-authority flow systems
- Queue age at the 75th percentile per lane is the earliest observable warning of approval-gate congestion because Little's Law implies that queue length and wait time growth co-move, and the P75 responds to developing congestion before the median or meanLittle (1961)Reinertsen (2009)
- Exception volume expressed as the four-week rolling Class 3 share of total intake is the most direct precursor to exception lane saturation and provides an estimated one to two weeks of warning before the Q3 circuit-breaker firesMitchell (2026)Mitchell (2026)
- DORA's change lead time, measured as a four-week moving average per lane, integrates approval-gate congestion and pipeline health into one observable and functions as a composite confirmation metric rather than a first-line signal; it is most useful when it rises in combination with earlier Tier 1 and Tier 2 signalsDORA (2024)Mitchell (2026)
- A deployment rework rate sustained above 20% of total deployments for two or more consecutive sprints signals that unplanned remediation work is consuming capacity that would otherwise be available for planned delivery, consistent with the SRE toil threshold at which reactive work becomes structurally damagingGoogle SRE Workbook (n.d.)DORA (2024)
- A persistent deployment-to-ticket gap exceeding 10% for three or more consecutive sprints signals systematic shadow-work growth, which predicts either escalating compliance risk from undisclosed changes or a future formal-queue surge when the informal backlog is regularised under audit pressureMitchell (2026)Mitchell (2026)
- Multi-window threshold design, requiring simultaneous breach of a short-window (one to two weeks) and a long-window (four weeks) check before an alert fires, reduces false positives from batch-arrival spikes while preserving detection speed for sustained congestion, applying the same alerting design principle demonstrated in SRE error budget burn-rate practiceGoogle SRE Workbook (n.d.)Reinertsen (2009)
- Rising exception share combined with a stable or declining rework rate signals deliberate misclassification of exception-path work as fast-path work, a behavioural coping response to governance friction that undermines routing integrity without triggering the standard exception volume alarmMitchell (2026)Mitchell (2026)
- Pre-agreed thresholds aligned to Q3 circuit-breaker conditions and Q5 control-regime shift points enable automatic regime adjustments without per-incident negotiation, following the same operational logic as the SRE error budget policy, which removed deliberation cost from the high-stakes decision of whether to continue feature deployment or switch to reliability-first operationsGoogle SRE Workbook (n.d.)Mitchell (2026)
Research Question
Which metrics best predict unsafe queue growth, rising delivery risk, or hidden demand accumulation in a split-authority delivery system, where "split-authority" means a context in which authority is divided among at least two stakeholder groups with independent veto or approval power?
Findings
Executive Summary
Five metric families, arranged into four warning tiers by causal distance from the approval constraint, constitute the minimum viable early-warning telemetry for split-authority delivery systems: queue age at the 75th percentile per lane, blocked-work ratio, exception volume as the four-week rolling Class 3 share of intake, lead time trend as a four-week moving average, and rework plus unplanned capacity share. A sixth proxy, the deployment-to-ticket gap, surfaces shadow-work growth, which signals governance-friction-driven compliance risk rather than throughput risk. The tiered structure maps directly to the routing circuit-breakers in Q3 and the control-regime shift conditions in Q5, enabling pre-agreed automatic responses that eliminate per-incident deliberation cost. Threshold values must be calibrated to each organisation's baseline; the proposed starting-point values are design heuristics grounded in Little's Law, DORA, and SRE burn-rate alerting precedent rather than empirically universal constants.
Key Findings
-
Queue age at the 75th percentile per lane is the earliest observable warning of approval-gate congestion because Little's Law implies that queue length and wait time growth co-move, and the P75 responds to developing congestion before the median or mean.
-
Exception volume expressed as the four-week rolling Class 3 share of total intake is the most direct precursor to exception lane saturation and provides an estimated one to two weeks of warning before the Q3 circuit-breaker fires.
-
DORA's change lead time, measured as a four-week moving average per lane, integrates approval-gate congestion and pipeline health into one observable and functions as a composite confirmation metric rather than a first-line signal; it is most useful when it rises in combination with earlier Tier 1 and Tier 2 signals.
-
A deployment rework rate sustained above 20% of total deployments for two or more consecutive sprints signals that unplanned remediation work is consuming capacity that would otherwise be available for planned delivery, consistent with the SRE toil threshold at which reactive work becomes structurally damaging.
-
A persistent deployment-to-ticket gap exceeding 10% for three or more consecutive sprints signals systematic shadow-work growth, which predicts either escalating compliance risk from undisclosed changes or a future formal-queue surge when the informal backlog is regularised under audit pressure.
-
Multi-window threshold design, requiring simultaneous breach of a short-window (one to two weeks) and a long-window (four weeks) check before an alert fires, reduces false positives from batch-arrival spikes while preserving detection speed for sustained congestion, applying the same alerting design principle demonstrated in SRE error budget burn-rate practice.
-
Rising exception share combined with a stable or declining rework rate signals deliberate misclassification of exception-path work as fast-path work, a behavioural coping response to governance friction that undermines routing integrity without triggering the standard exception volume alarm.
-
Pre-agreed thresholds aligned to Q3 circuit-breaker conditions and Q5 control-regime shift points enable automatic regime adjustments without per-incident negotiation, following the same operational logic as the SRE error budget policy, which removed deliberation cost from the high-stakes decision of whether to continue feature deployment or switch to reliability-first operations.
-
The specific starting-point threshold values proposed (P75 age rising across two consecutive measurements; exception share rising five percentage points over four weeks; 20% unplanned capacity share for two sprints; 10% deployment-to-ticket gap for three sprints) are design heuristics that must be calibrated against each organisation's six-month baseline before operational deployment.
Assumptions
-
The exception review function is the binding flow constraint in the split-authority system; if a testing environment, vendor dependency, or staffing shortage is the actual constraint, the indicator bundle must be re-anchored to that constraint, and Q1 constraint identification must be completed first.
-
Emergency changes processed under a registered emergency change practice are excluded from the shadow-work proxy; if no such register exists, the proxy overstates shadow-work growth, reducing its specificity.
-
The proposed threshold values are calibration starting points; each organisation should collect six months of baseline data before operationalising the warning bundle, since natural variation in queue age and exception volume differs across delivery contexts.
Analysis
Monitoring queue dynamics through a tiered indicator bundle requires understanding the causal chain: intake volume and mix drive approval-gate service capacity toward queue age and WIP accumulation, then toward lead time growth, rework rate increase, and, if unchecked, shadow-work proliferation as teams adopt coping strategies. Non-linear queueing dynamics imply that wait time growth accelerates as utilisation approaches capacity, which means early indicators (Tier 1 and Tier 2) provide proportionally more response time than late indicators (Tier 3 and Tier 4). This asymmetry justifies the investment in queue-age and exception-volume monitoring even when those metrics are harder to automate than lead time, which is directly available from most delivery pipeline tools.
Shadow-work proxy signals require separate treatment in the analysis because a rising proxy indicates a different failure mode: governance-legitimacy collapse rather than throughput collapse. When shadow-work growth rises, the response is to reduce intake friction (make the formal process easier than the informal path) rather than to tighten routing controls. Tightening controls in response to shadow-work growth typically accelerates the coping behaviour rather than reversing it.
Key Finding 7 (misclassification signal) is a cross-indicator inference: neither rising exception share alone nor declining rework rate alone constitutes the signal; only the combination does. This makes it harder to automate but important to include, because misclassification erodes the integrity of the demand-segmentation model without producing an obvious throughput signal until field defects materialise.
Risks, Gaps, and Uncertainties
-
No empirical study directly validates threshold values for approval queue leading indicators in regulated split-authority governance systems; all proposed values are heuristics. The evidence base for threshold design comes from analogous domains (SRE, product development flow) rather than from direct observation of approval-gate congestion events.
-
DORA survey data primarily reflects software development team experience and may not fully represent regulated enterprise environments where compliance and risk functions with different staffing, skill, and authority structures form part of the approval chain. The completed HITL capacity thresholds study confirms that regulated enterprises calibrate review-queue thresholds locally using supervisory guidance (such as HKMA requirements for clear internal timelines) rather than universal benchmarks, reinforcing that DORA-derived starting values require regulatory-context adjustment before operational deployment.
-
The shadow-work proxy depends on integrated deployment tracking and ticket systems; organisations with manual or fragmented toolchains will have difficulty operationalising it without additional instrumentation investment. [assumption; justification: most modern cloud-native delivery pipelines have automated deployment tracking; legacy environments may not]
-
The constraint assumption (that exception review is the binding constraint) has not been validated as a standalone Q1 research item; the indicator bundle is correctly calibrated only for approval-gate-constrained systems.
Open Questions
-
What threshold values for approval queue indicators have been observed to predict queue collapse events in regulated enterprise environments? A dedicated empirical study collecting baseline metrics and observing saturation events would provide the calibration evidence this item lacks.
-
Can the misclassification signal be operationalised more reliably by adding a periodic classification-accuracy audit (sampling recent Class 3 items against the Q2 boundary tests) rather than relying solely on the cross-indicator combination?
-
How should the indicator bundle be adapted when the binding constraint is an external regulatory review body rather than an internal approval gate? External review service-time distributions are different, and Little's Law calibration would require different baseline assumptions.
-
Does the multi-window alerting design translate from SRE error budgets (continuous event streams) to approval queues (discrete batch events) without modification, or does the lower event frequency require different window lengths?
sources
- [x] DORA (2024) Software delivery metrics: throughput and instability
- [x] Google SRE Workbook: Alerting on Service Level Objectives
- [x] Google SRE Workbook: Eliminating Toil
- [x] Little, J.D.C. (1961) A proof for the queuing formula, Operations Research
- [x] Reinertsen, D.G. (2009) The Principles of Product Development Flow, Celeritas Publishing
- [x] TOC Institute (2024) Five Focusing Steps
- [x] Mitchell (2026) Operating model synthesis for split-authority delivery systems (P1)
- [x] Mitchell (2026) Q2: Demand segmentation for fast-path vs controlled-path flow
- [x] Mitchell (2026) Q3: Routing design that isolates exceptions from routine flow
- [x] Mitchell (2026) Q4: Decision rights that should move closer to execution
- [x] Mitchell (2026) Q5: Control model for the best throughput-risk trade-off
- [x] Mitchell (2026) Backpressure infrastructure and the Theory of Constraints
- [x] Mitchell (2026) Internal governance controls: effectiveness conditions in regulated enterprises
- [x] Mitchell (2026) Governance failure mechanisms: bureaucracy and circumvention
- [x] Mitchell (2026) Variance control comparison across delivery modes