Q5: Control model for the best throughput-risk trade-off

2026-05-29 · governance-policy organisational-design cost-performance enterprise-adoption regulatory-compliance · medium · source → · wiki →
key claims
  1. Post-hoc review (a detective control operating after the event) is the correct and sufficient control for Class 1 work because the three Q2 boundary tests already provide the ex-ante governance: a passing template test confirms pattern adherence, a passing recovery test confirms reversibility, and a passing blast radius test confirms contained impact, leaving post-hoc review to confirm adherence and update the template catalogueMitchell (2026)COSO (n.d.)PeopleCert (n.d.)
  2. Bounded delegation with guardrails (a hybrid governance structure in TCE terms) is the correct control for Class 2 work, in which the three Q4 parameters (cost ceiling, blast radius limit, approved technology catalog) and four escalation triggers (catalog deviation, ceiling breach, blast radius overflow, external commitment) together constitute the control, replacing per-decision pre-approval while preserving governance outcomes for the boundary cases that matterMitchell (2026)Williamson (1979)
  3. Pre-approval is the correct control for Class 3 work (novel, irreversible, or system-wide blast radius), where high uncertainty and high consequence of contracting failure justify the throughput cost of ex-ante review, consistent with TCE hierarchical governance and ITIL 4's Emergency Change Advisory Board (ECAB) model for emergency and novel-consequence changesWilliamson (1979)PeopleCert (n.d.)
  4. Decision velocity is an override criterion: time-critical decisions (incident response, emergency mitigation) cannot use pre-approval regardless of class assignment because approval latency is structurally incompatible with the five-minute response window for services targeting four nines (99.99%) availability, and must instead use pre-delegated Incident Command System (ICS) authority with post-hoc reviewSre (n.d.)Sre (n.d.)
  5. Detection velocity sets the boundary condition for post-hoc review acceptability: post-hoc review is risk-acceptable only when automated monitoring can detect failure before material harm accumulates, as demonstrated by the SRE error budget burn rate as a near-real-time detection mechanism for Class 1 deployment failuresGoogle SRE Book (n.d.)Google SRE Workbook (n.d.)
  6. The SRE error budget policy operationalises a trigger-based shift from bounded delegation (normal operations, below Service Level Objective (SLO) budget threshold) to pre-approval (deployment freeze when error budget is exhausted), providing a concrete real-world example of automatic control-pattern escalation governed by a measurable threshold rather than a subjective judgmentGoogle SRE Workbook (n.d.)
  7. Applying pre-approval to Class 1 work generates approval-queue latency without proportionate risk reduction, because the work is reversible and contained by definition; this pattern constitutes a coercive formalisation in Adler and Borys (1996) terms and predictably generates mis-classification and workaround behaviour as documented in the governance failure mechanisms item in this corpusBorys (1996)Github (n.d.)Mitchell (2026)
  8. Two observable conditions trigger temporary tightening of the control regime by one level: exception volume ratio above a calibrated threshold (from the Q3 circuit-breaker model), and error budget exhaustion (from the SRE error budget policy); both conditions indicate that the current class boundary tests are failing to contain risk at their defined levelsMitchell (2026)Google SRE Workbook (n.d.)

Research Question

When should the system use pre-approval, bounded delegation with guardrails, or post-hoc review and exception escalation?

Findings

Executive Summary

Post-hoc review is the correct and sufficient control for Class 1 (fast-path) work, bounded delegation with guardrails is the correct control for Class 2 (standard-path) work, and pre-approval is the correct control for Class 3 (exception-path) work, with class assignment determined by the three Q2 boundary tests (template, recovery, blast radius). This mapping is independently supported by Transaction Cost Economics (TCE) discriminating alignment (governance intensity should match transaction hazard), the Committee of Sponsoring Organizations of the Treadway Commission (COSO) preventive-detective distinction, Information Technology Infrastructure Library version 4 (ITIL 4) Change Enablement practice, and the Site Reliability Engineering (SRE) error budget policy as a concrete implementation of trigger-based regime shifts. Applying pre-approval to Class 1 work imposes approval-queue latency without proportionate risk reduction, because the boundary tests that qualify Class 1 work already confirm reversibility, contained blast radius, and pattern adherence. Two observable triggers (exception volume ratio above threshold and error budget exhausted) shift the control regime one level toward pre-approval temporarily; two observable conditions (exception volume ratio below threshold for a sustained window and post-hoc review confirming sustained Class 1 compliance) permit shifting one level toward post-hoc review.

Key Findings

  1. Post-hoc review (a detective control operating after the event) is the correct and sufficient control for Class 1 work because the three Q2 boundary tests already provide the ex-ante governance: a passing template test confirms pattern adherence, a passing recovery test confirms reversibility, and a passing blast radius test confirms contained impact, leaving post-hoc review to confirm adherence and update the template catalogue.

  2. Bounded delegation with guardrails (a hybrid governance structure in TCE terms) is the correct control for Class 2 work, in which the three Q4 parameters (cost ceiling, blast radius limit, approved technology catalog) and four escalation triggers (catalog deviation, ceiling breach, blast radius overflow, external commitment) together constitute the control, replacing per-decision pre-approval while preserving governance outcomes for the boundary cases that matter.

  3. Pre-approval is the correct control for Class 3 work (novel, irreversible, or system-wide blast radius), where high uncertainty and high consequence of contracting failure justify the throughput cost of ex-ante review, consistent with TCE hierarchical governance and ITIL 4's Emergency Change Advisory Board (ECAB) model for emergency and novel-consequence changes.

  4. Decision velocity is an override criterion: time-critical decisions (incident response, emergency mitigation) cannot use pre-approval regardless of class assignment because approval latency is structurally incompatible with the five-minute response window for services targeting four nines (99.99%) availability, and must instead use pre-delegated Incident Command System (ICS) authority with post-hoc review.

  5. Detection velocity sets the boundary condition for post-hoc review acceptability: post-hoc review is risk-acceptable only when automated monitoring can detect failure before material harm accumulates, as demonstrated by the SRE error budget burn rate as a near-real-time detection mechanism for Class 1 deployment failures.

  6. The SRE error budget policy operationalises a trigger-based shift from bounded delegation (normal operations, below Service Level Objective (SLO) budget threshold) to pre-approval (deployment freeze when error budget is exhausted), providing a concrete real-world example of automatic control-pattern escalation governed by a measurable threshold rather than a subjective judgment.

  7. Applying pre-approval to Class 1 work generates approval-queue latency without proportionate risk reduction, because the work is reversible and contained by definition; this pattern constitutes a coercive formalisation in Adler and Borys (1996) terms and predictably generates mis-classification and workaround behaviour as documented in the governance failure mechanisms item in this corpus.

  8. Two observable conditions trigger temporary tightening of the control regime by one level: exception volume ratio above a calibrated threshold (from the Q3 circuit-breaker model), and error budget exhaustion (from the SRE error budget policy); both conditions indicate that the current class boundary tests are failing to contain risk at their defined levels.

  9. Two observable conditions permit relaxing the control regime by one level: exception volume ratio below threshold sustained over a calibrated observation window, and post-hoc review over the same window confirming that Class 1 work consistently meets its boundary conditions without in-flight escalation; the SRE model explicitly permits lifting the deployment freeze once SLO compliance is restored.

  10. The Human-in-the-Loop (HITL) capacity constraint provides an independent upper bound on how much work can be routed through pre-approval before meaningful review collapses into rubber-stamping (approval-by-exception without genuine challenge), and the conditional control-selection matrix must therefore be designed so that Class 3 volume stays within the genuine review capacity of the pre-approval authority.

Assumptions

  1. Automated monitoring provides near-real-time failure detection for Class 1 work. Justification: SRE and DevOps Research and Assessment (DORA) both treat automated monitoring as a prerequisite for fast-path deployment; without it, detection velocity may be insufficient for post-hoc review to be risk-acceptable.

  2. Governance parameters (cost ceiling, blast radius limit, approved catalog) for bounded delegation are maintained on a regular cadence by the central function; stale parameters undermine the bounded delegation model and force fallback to pre-approval. Justification: the Q4 item identified parameter maintenance as the primary operational assumption for bounded delegation.

  3. Exception volume ratio thresholds and error budget thresholds are organisation-specific and must be calibrated empirically; the evidence provides the mechanism but not universal threshold values. Justification: the SRE error budget model explicitly leaves SLO values to be set per service; Q3 similarly leaves circuit-breaker thresholds to empirical calibration.

Analysis

The five selection criteria (reversibility, blast radius, standardisation, decision velocity, detection velocity) are not independent: reversibility and blast radius together determine whether the failure mode can be corrected before material harm accumulates, and standardisation determines whether the failure mode is known in advance. Decision velocity and detection velocity together determine whether ex-ante or ex-post review can operationally deliver governance value in the available time.

Uniform pre-approval is the primary rival to the conditional matrix, and has historically been treated as a compliance default in regulated enterprises. Multiple independent sources converge against uniform pre-approval: the P1 item identified governance-generated queueing as the dominant flow constraint, the HITL capacity thresholds item showed that pre-approval at high volume degrades to rubber-stamping (approval-by-exception without genuine challenge), and the SRE model demonstrates that even regulated-context governance can use bounded delegation as the default with pre-approval reserved for budget-exhaustion events.

The one scenario where uniform pre-approval is defensible is when the organisation cannot maintain automated monitoring (making detection velocity insufficient for post-hoc review) and cannot maintain the governance parameters that make bounded delegation reliable; in that case, the absence of the enabling conditions for both alternative patterns forces the default back to pre-approval despite its throughput cost, consistent with the competence-prerequisite scenario identified in Q4.

The Basel Committee on Banking Supervision (BCBS) 328 proportionality requirement provides the regulatory floor: control intensity must match risk profile. Uniform pre-approval exceeds the required proportionality for Class 1 work; uniform post-hoc review falls below the required proportionality for Class 3 work. The conditional matrix is structurally consistent with BCBS 328 because each demand class carries a named control pattern, a named authority (from Q4), and trigger-defined escalation paths.

The behavioural risk is that teams will mis-classify Class 2 or 3 work as Class 1 to access the post-hoc review lane. The Q2 boundary tests (observable, self-evident, auditable) reduce this incentive by making classification legible. The Q3 Work in Progress (WIP) limit on the exception lane removes the incentive to inflate Class 3 status for priority service. The conditional regime-shift mechanism adds a system-level deterrent: sustained mis-classification raises exception volume, triggering the tighten condition that shifts all lanes toward pre-approval, which penalises compliant teams alongside mis-classifying teams and therefore creates collective pressure to maintain accurate classification.

Risks, Gaps, and Uncertainties

Open Questions

  1. How should the tighten and relax observation windows be sized for teams with varying delivery cadence?
  2. What monitoring design detects "silent failure" scenarios (where Class 1 failures are slow-developing and not captured by automated monitoring) without recreating pre-approval latency?
  3. At what Class 3 volume does the pre-approval mechanism exceed the genuine review capacity of the authority function, triggering the HITL rubber-stamp failure mode?

Output

sources


cites
cites Q2: Demand segmentation for fast-path vs controlled-path flow
cites Q3: Routing design that isolates exceptions from routine flow
cites Q4: Decision rights that should move closer to execution
cites Operating model synthesis for split-authority delivery systems
cites Conditions under which internal governance controls minimise coordination costs in regulated enterprises
cites Failure mechanisms of internal governance controls: bureaucratic inefficiency and informal circumvention in regulated enterprises
related (frontmatter)
related At what threshold does Human-in-the-Loop (HITL) oversight in bank compliance operations stop being a meaningful challenge function and become routine acceptance of automated outputs?
related How should Artificial Intelligence (AI) and low-code use cases be classified into risk tiers, and how should governance controls vary across those tiers?
related Regulatory and standards preconditions for deployment of Artificial Intelligence (AI) systems that can take multi-step actions: does incomplete access control and data governance constitute a control failure?
version history
versiondatecommitsummary
1.02026-05-314f19547Initial completion

Connected items

Loading…

View full knowledge graph →