Governance and operating models for safe-to-fail experimentation in regulated…
Governance and operating models for safe-to-fail experimentation in regulated industries
- Regulator-run financial sandboxes are built around a formal legal or administrative mandate, a fixed testing period, transparent selection criteria, and defined exit conditions, distinguishing them from informal regulatory forbearanceLauer (2017)Financial (n.d.)
- Entry into the United Kingdom's Financial Conduct Authority sandbox between 2014 and 2019 is associated with a 15% increase in the average amount of capital raised and a 50% increase in the probability of raising capital, with larger effects for smaller and younger firmsMerrouche (2023)
- Of 146 applications received across the first two cohorts of the Financial Conduct Authority sandbox, 50 firms were accepted into structured testing, and the regulator reports early evidence of reduced time and cost to market for participantsFCA (2017)
- Health-sector regulatory sandboxes such as the Medicines and Healthcare products Regulatory Agency's AI Airlock target named, live regulatory questions such as synthetic-data validation and post-market surveillance rather than open-ended exploration, and feed generated evidence forward into procurement and future regulatory decisionsMedicines (n.d.)Www (n.d.)
- The United States Food and Drug Administration's five-year Digital Health Software Precertification pilot concluded that scaling its organisation-level oversight model beyond the pilot would require new statutory authority from Congress, demonstrating a governance-structure ceiling that sits outside the sandbox's own designFDA (n.d.)
- Academic critique of the regulatory sandbox model argues that sandboxes achieve safe experimentation partly by temporarily relaxing consumer-protection and prudential requirements rather than through governance structure alone, and predicts that competition among jurisdictions for fintech business will push sandbox rules toward looser boundary conditions over timeAllen (2019)
- Embedded-control team governance, as defined by the Institute of Internal Auditors' Three Lines Model, distributes risk accountability to first-line operational management while second-line risk and compliance functions provide oversight and challenge rather than holding a sequential go/kill veto typical of a stage-gated reviewInstitute (n.d.)
- None of the regulator, standards-body, or academic sources consulted for this item, including the Financial Conduct Authority sandbox documentation, the Bank for International Settlements working paper, and the Institute of Internal Auditors Three Lines Model, measures the relationship between the structure of an internal experiment-tracking and prioritisation mechanism and experiment volume, pattern quality, compliance incidents, or adoption specifically inside a regulated organisationFinancial (n.d.)Merrouche (2023)Institute (n.d.)
Research Question
In highly regulated industries such as financial services, healthcare, and pharmaceuticals, how do organisations design governance structures, team models, and operating practices that enable safe-to-fail probing experiments in the Complex domain of the Cynefin framework, and what empirical relationship exists between the degree of structure in experiment tracking and prioritisation mechanisms and outcomes including (a) volume and diversity of experiments conducted, (b) quality and actionability of learned patterns, (c) regulatory compliance and risk incidents, and (d) organisational adoption of emergent innovations?
Findings
(Populated from §6 Synthesis above.)
Executive Summary
Regulated organisations use four distinct governance archetypes for safe-to-fail probing, regulator-run financial sandboxes, health-sector regulatory sandboxes, organisation-level precertification, and embedded-control teams, but no source located in this investigation directly measures how the structure of an internal experiment-tracking mechanism affects experiment volume, pattern quality, compliance incidents, or adoption inside a regulated firm. This item's evidence base contains no quantified relationship measuring internal tracking-mechanism structure directly, only one measuring a governance-structure choice: sandbox participation. Entering a regulator-run sandbox, a governance-structure choice rather than an internal tracking-mechanism choice, is associated with a 15% increase in capital raised and a 50% increase in the probability of raising capital at all for participating fintech firms. Regulated safe-to-fail programmes bound their probes with legal mandates, fixed testing windows, participant caps, and continuous supervisory oversight rather than the informal guardrails of Cynefin's original safe-to-fail probe design. The relationship between governance structure and outcomes is non-linear in at least one documented direction: the FDA's five-year Pre-Cert pilot generated useful learning but concluded that scaling its oversight model would exceed existing statutory authority, showing that a well-run safe-to-fail experiment can still fail to convert into a scaled operating model when the binding constraint sits above the sandbox's internal design. Only one academic critique of the sandbox archetype's safety claim was located in this item's evidence base. It argues that sandboxes achieve safe experimentation partly by temporarily suspending consumer-protection requirements rather than purely through governance structure, and predicts this trade-off will loosen as jurisdictions compete for fintech business.
Key Findings
- Regulator-run financial sandboxes are built around a formal legal or administrative mandate, a fixed testing period, transparent selection criteria, and defined exit conditions, distinguishing them from informal regulatory forbearance.
- Entry into the United Kingdom's Financial Conduct Authority sandbox between 2014 and 2019 is associated with a 15% increase in the average amount of capital raised and a 50% increase in the probability of raising capital, with larger effects for smaller and younger firms.
- Of 146 applications received across the first two cohorts of the Financial Conduct Authority sandbox, 50 firms were accepted into structured testing, and the regulator reports early evidence of reduced time and cost to market for participants.
- Health-sector regulatory sandboxes such as the Medicines and Healthcare products Regulatory Agency's AI Airlock target named, live regulatory questions such as synthetic-data validation and post-market surveillance rather than open-ended exploration, and feed generated evidence forward into procurement and future regulatory decisions.
- The United States Food and Drug Administration's five-year Digital Health Software Precertification pilot concluded that scaling its organisation-level oversight model beyond the pilot would require new statutory authority from Congress, demonstrating a governance-structure ceiling that sits outside the sandbox's own design.
- Academic critique of the regulatory sandbox model argues that sandboxes achieve safe experimentation partly by temporarily relaxing consumer-protection and prudential requirements rather than through governance structure alone, and predicts that competition among jurisdictions for fintech business will push sandbox rules toward looser boundary conditions over time.
- Embedded-control team governance, as defined by the Institute of Internal Auditors' Three Lines Model, distributes risk accountability to first-line operational management while second-line risk and compliance functions provide oversight and challenge rather than holding a sequential go/kill veto typical of a stage-gated review.
- None of the regulator, standards-body, or academic sources consulted for this item, including the Financial Conduct Authority sandbox documentation, the Bank for International Settlements working paper, and the Institute of Internal Auditors Three Lines Model, measures the relationship between the structure of an internal experiment-tracking and prioritisation mechanism and experiment volume, pattern quality, compliance incidents, or adoption specifically inside a regulated organisation.
- General new-product-development literature attributes a large share of documented innovation project failures to errors in upfront planning and design rather than execution or marketization errors, a pattern consistent with, though not proof of, under-structured experiment tracking contributing to failure. ([inference]; low confidence; source: abstract of Coccia (2023) accessed via secondary aggregation of Coccia (2023) New Perspectives in Innovation Failure Analysis
- Technology-sector experimentation literature argues that qualitative feedback and simple experiment counts cannot reliably show whether a change to tracking or review process improved experimentation quality, and instead proposes applying controlled testing to the experimentation process itself; this evidence originates outside regulated industries and is used here as an analogy rather than direct regulated-sector evidence.
Assumptions
This item assumes that findings from technology-sector experimentation platforms transfer as directional evidence, not as regulated-sector proof, to regulated-industry safe-to-fail probing. This assumption is justified because no source located studies internal experiment-tracking mechanism structure specifically inside a regulated organisation, making the technology-sector literature the closest available body of evidence on how tracking structure interacts with learning quality, even though the risk and review context differs substantially.
This item assumes that the FCA sandbox's documented capital-raising effect and the FDA Pre-Cert pilot's documented scaling ceiling are both representative of governance-structure effects in their respective sectors rather than idiosyncratic to the specific programmes studied. This assumption is justified because each is the most rigorously evaluated example located in its sector, an econometric matched-control study for the FCA case and a five-year government pilot with a public final conclusion for the FDA case, but neither has an independent replication in a second jurisdiction within the evidence base gathered here.
Analysis
The evidence separates cleanly into two tiers of directness. The first tier, governance-structure choices such as whether to enter a regulator-run sandbox at all, has one rigorously quantified outcome relationship, because that comparison has a natural control group of similar firms that did not enter the sandbox. The second tier, internal tracking-mechanism structure once inside a regulated experimentation programme, has no equivalent natural experiment in the located evidence base, because firms rarely publish comparisons of their own internal backlog or portfolio-board practices, and regulators do not require or collect that level of internal process detail.
A plausible rival explanation for the capital-raising effect is that firms selected into the sandbox were already stronger candidates before entry, meaning the sandbox itself adds a credibility signal rather than causing genuine quality improvement. The BIS paper addresses this by using a matched-control design and by showing the effect concentrates among firms facing the largest prior informational disadvantage, which weakens the pure-selection explanation without eliminating it, since selection into any sandbox cohort is itself non-random.
The FDA Pre-Cert and MHRA AI Airlock cases represent two different resolutions of the same underlying tension between speed of learning and depth of statutory change. The FDA scoped its pilot toward a permanent shift in the oversight model itself and found that shift blocked by statute, whereas the MHRA scoped its sandbox toward narrower, per-topic regulatory questions intended to inform incremental future guidance rather than a wholesale precertification regime, which may explain why the MHRA programme continued into a second, expanded phase while the FDA pilot concluded without a scaled successor.
An alternative remedy to changing the governance model, adding review staff or strengthening model-quality gates instead of redesigning the operating model, is not directly evidenced as tried or rejected in either the FDA or MHRA case; both programmes moved toward operating-model change rather than toward scaling review headcount, but no source located explains why headcount scaling was not the chosen alternative.
Risks, Gaps, and Uncertainties
The central research question, the empirical relationship between tracking-and-prioritisation mechanism structure and the four named outcomes, remains substantially unanswered by direct regulated-sector evidence; every source located that discusses tracking-mechanism structure and outcomes is drawn from technology-sector experimentation or general new-product-development literature rather than regulated-industry safe-to-fail programmes specifically.
The Coccia (2023) source could not be accessed beyond its abstract: both the ScienceDirect article page and a ResearchGate-hosted preprint mirror returned HTTP 403 errors when fetched directly in this session, so all claims drawn from it are held at inference confidence rather than fact.
The MHRA AI Airlock pilot's named case-study participants and topics could not be verified against the primary PDF report text, which the available fetch tool returned only as unparsed binary content; the claim rests on a secondary aggregation of the report and is held at inference confidence pending direct textual confirmation.
The Cofie (2024) practitioner source listed in the seeded Sources is no longer accessible: the original URL now redirects to a page with no article content, and no working alternative copy was located during this session; it is retained in the Sources list as identified-but-not-consulted and contributes no claims to this item.
No source located provides a like-for-like, cross-jurisdiction comparison of sandbox boundary-condition strictness over time, so Allen's prediction of a regulatory race-to-the-bottom in sandbox design remains a single-author analytical argument rather than an empirically confirmed trend.
The BIS working paper's capital-raising findings are specific to the FCA's sandbox and the 2014-2019 period; no equivalent econometric evaluation of the Monetary Authority of Singapore, Hong Kong Monetary Authority, or other national sandboxes was located, so the generalisability of the 15% and 50% effect sizes to other regulatory sandbox designs is unconfirmed.
Open Questions
Does the internal structure of an experiment-tracking and prioritisation mechanism measurably affect experiment volume, pattern quality, compliance incidents, or adoption inside a regulated organisation, independent of whether the organisation also participates in an external regulator-run sandbox?
Has any national financial or health regulator published a like-for-like, multi-year comparison of sandbox outcomes across two or more jurisdictions that would allow Allen's race-to-the-bottom prediction to be tested empirically rather than argued analytically?
What accounts for the different scaling trajectories of the FDA's Pre-Cert pilot, which concluded without a scaled successor, and the MHRA's AI Airlock, which continued into an expanded second phase: is it the narrower scope of the AI Airlock's per-topic questions, a difference in statutory flexibility between the two regulators, or some other factor not captured in the sources reviewed here?
sources
- [x] Cynefin.io - Cynefin - official overview of the framework and its Complex-domain framing
- [x] Cynefin.io - Safe to fail probes - official explanation of safe-to-fail probe design and use
- [x] Financial Conduct Authority (FCA) Regulatory Sandbox - official United Kingdom sandbox model for controlled experimentation
- [x] FCA (2017) Regulatory sandbox lessons learned report - first-cohort outcome data on applications, acceptance, and impact
- [x] Organisation for Economic Co-operation and Development (OECD) Regulatory Sandbox Toolkit, Digital Object Identifier (DOI) - public toolkit for designing and governing regulatory sandboxes (original oecd.org URL blocked to automated fetch in this session; DOI resolves to the same official publication)
- [x] Müller (2024) Meta-experiments: Improving experimentation through experimentation - experimentation-on-experimentation evidence relevant to tracking and prioritisation design
- [x] Coccia (2023) New Perspectives in Innovation Failure Analysis - taxonomy of innovation failure and risk-reduction strategies
- [ ] Cofie (2024) Innovating in a straitjacket; a guide to navigating innovation in highly regulated industries - practitioner framing across banking, insurance, and healthcare; page no longer resolves to article content, identified but not consulted
- [x] Cornelli, Doerr, Gambacorta, Merrouche (2023) Regulatory sandboxes and fintech funding: evidence from the UK, BIS Working Paper 901 - Bank for International Settlements (BIS) econometric evaluation of sandbox effects on capital raised, survival, and patenting
- [x] Jenik and Lauer (2017) Regulatory Sandboxes and Financial Inclusion, Consultative Group to Assist the Poor (CGAP) Working Paper - design elements, constraints, and limits of sandboxes in emerging markets and developing economies (EMDEs)
- [x] Allen (2019) Sandbox Boundaries, Vanderbilt Journal of Entertainment and Technology Law 22(2) - critical academic analysis of consumer-protection trade-offs and cross-border information-sharing limits in sandboxes
- [x] Medicines and Healthcare products Regulatory Agency (MHRA) - AI Airlock Sandbox Pilot Programme Report - United Kingdom healthcare AI regulatory sandbox pilot findings
- [x] GOV.UK - Pioneering AI health innovations regulatory sandbox launched - National Health Service (NHS) and MHRA sandbox expansion announcement
- [x] FDA Digital Health Software Precertification (Pre-Cert) Pilot Program - United States Food and Drug Administration (FDA) organisation-level oversight pilot and its scaling limits
- [x] Institute of Internal Auditors (IIA) - The IIA's Three Lines Model - authoritative definition of the three-lines governance architecture used in embedded-control team designs
- [x] Kohavi, Tang, and Xu (2020) Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing - experimentation-platform maturity model and tracking-mechanism evidence, used as a technology-sector analogy
- [x] Cooper (Stage-Gate International) The Stage-Gate Model: An Overview - structured new-product-development portfolio-tracking model, used as a general innovation-management analogy