Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic…
Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic Systems Under Distribution Shift
- Accessible summaries of Friedman's methodology treat predictive fruitfulness, not realism of assumptions, as the decisive test of a theory, which makes instrumentalism an epistemic rule for accepting models that predict well without requiring them to explain the underlying mechanismFriedman (1953)Hausman (2018)Essays (n.d.)
- Goodhart's Law shows that once an observed regularity is pressed into service as a control target, optimisation pressure can change the underlying process and collapse the regularity, so score-maximisation itself becomes a source of model failure rather than proof of model adequacyGoodhart (1975)Australia (1990)Goodhart (n.d.)
- Pearl's causal hierarchy and later invariance work jointly imply that observationally successful non-causal predictors can become badly wrong under intervention or environmental change, because only causal or invariant structure is expected to travel across such shiftsPearl (2008)Pearl (2009)Meinshausen (2015)Buhlmann (2018)
- Sugihara et al. show that nonlinear dynamic systems can exhibit mirage correlations and require tools stronger than correlation or one-step predictability to detect causation, which means short-run predictive success in such systems can conceal causal misidentificationSugihara et al. (2012)
- Economic forecasting under structural instability fails through regime brittleness, because crisis periods such as 2007-08 invalidate the assumption that average historical performance is still informative, and instability-aware evaluation must replace simple retrospective scorekeepingRossi (2021)Perron (2018)Settlements (2008)
- Coronavirus Disease 2019 (COVID-19) case forecasting exposed assumption lock-in, because many official-hub models failed to beat simple baselines and were built around continuation assumptions about interventions or behavior that became unreliable as policy, reporting, and variant conditions changedShah et al. (2024)Nixon et al. (2022)
- Recommendation systems exhibit silent quality decay under preference drift, because yesterday's successful correlations degrade in non-stationary user environments unless the system explicitly models drift, reweights evidence, or equips operators to interveneHinder et al. (2024)Pulungan (2025)
- The operational value of explanation is therefore that it makes diagnosis and repair more directed, because mechanistic or invariant accounts narrow what should remain stable, whereas instrumentalist systems rely more heavily on continual monitoring, recalibration, and governance after failure signals appearResearch (n.d.)Research (n.d.)Buhlmann (2018)Hinder et al. (2024)
Research Question
What are the operational failure modes of an epistemic framework that prioritises instrumentalism, treating predictive performance as the primary criterion, over explanatory reach when applied to complex dynamic systems undergoing distribution shift, changes in the data-generating environment or relationship structure?
Findings
Executive Summary
An instrumentalist modeling stance, which treats predictive fruitfulness as sufficient for accepting a model without requiring realistic assumptions or mechanistic explanation, fails in complex dynamic systems because predictive success without causal or mechanistic structure does not remain reliable when interventions, structural breaks, or adaptive gaming change the environment.
The recurring operational failure modes are metric-target deformation, regime brittleness, causal blindness, assumption lock-in, and silent quality decay.
The strongest support for that conclusion comes from causal hierarchy and invariance research, which shows that association-level success does not license confidence about interventions or shifted environments.
The case studies do not show that every predictive model fails under shift, but they do show that prediction-only success pushes more operational work into monitoring and recalibration because the model does not specify what should remain stable.
Key Findings
- Accessible summaries of Friedman's methodology treat predictive fruitfulness, not realism of assumptions, as the decisive test of a theory, which makes instrumentalism an epistemic rule for accepting models that predict well without requiring them to explain the underlying mechanism.
- Goodhart's Law shows that once an observed regularity is pressed into service as a control target, optimisation pressure can change the underlying process and collapse the regularity, so score-maximisation itself becomes a source of model failure rather than proof of model adequacy.
- Pearl's causal hierarchy and later invariance work jointly imply that observationally successful non-causal predictors can become badly wrong under intervention or environmental change, because only causal or invariant structure is expected to travel across such shifts.
- Sugihara et al. show that nonlinear dynamic systems can exhibit mirage correlations and require tools stronger than correlation or one-step predictability to detect causation, which means short-run predictive success in such systems can conceal causal misidentification.
- Economic forecasting under structural instability fails through regime brittleness, because crisis periods such as 2007-08 invalidate the assumption that average historical performance is still informative, and instability-aware evaluation must replace simple retrospective scorekeeping.
- Coronavirus Disease 2019 (COVID-19) case forecasting exposed assumption lock-in, because many official-hub models failed to beat simple baselines and were built around continuation assumptions about interventions or behavior that became unreliable as policy, reporting, and variant conditions changed.
- Recommendation systems exhibit silent quality decay under preference drift, because yesterday's successful correlations degrade in non-stationary user environments unless the system explicitly models drift, reweights evidence, or equips operators to intervene.
- The operational value of explanation is therefore that it makes diagnosis and repair more directed, because mechanistic or invariant accounts narrow what should remain stable, whereas instrumentalist systems rely more heavily on continual monitoring, recalibration, and governance after failure signals appear.
Assumptions
- The accessible secondary summaries of Friedman's essay and Goodhart's formulation are sufficient to characterise the methodological stance at issue because the original works are uniquely identified and the summaries track the relevant methodological claims closely enough for this item's level of analysis.
- The selected economic, epidemiological, and recommendation-system cases are representative enough to illustrate recurring operational failure classes without claiming identical proximal causes in every domain.
Analysis
Instrumentalism and explanatory evaluation fail differently under stress. A predictive model can look successful on one regime because it compresses observed regularities, but that success does not reveal whether the regularity is causal, merely correlative, or already being distorted by target-seeking behavior.
The strongest theoretical reason to expect failure under distribution shift comes from the causal hierarchy and invariance literature, not from any single case study. Those sources say that intervention robustness requires information above association-level fit, which supports the inference that distribution shift exposes exactly what instrumentalism declines to model.
A plausible rival explanation is that the observed failures came mainly from poor data quality or weak operations rather than from the epistemic stance of instrumentalism itself. That rival explanation is partly correct for COVID-19 and crisis forecasting, but it is incomplete because even perfect observational data do not answer intervention or counterfactual questions unless the model represents causal structure or stable invariants.
Another rival explanation is that continuous retraining, ensembles, or drift-aware adaptation make explanation unnecessary. Those remedies help, but they move the operational burden into ongoing monitoring and repair, which means they mitigate failure without showing that score-first modeling has captured the mechanism.
Risks, Gaps, and Uncertainties
- The accessible evidence for Friedman's and Goodhart's original texts is weaker than the evidence for Pearl, Peters, Buhlmann, COVID-19 forecasting, and concept drift, because the seeded primary landing pages were unavailable and the argument depends partly on high-quality secondary summaries.
- The economic case evidence is strongest on instability and evaluation method, not on a single universally agreed post-mortem that names instrumentalism as the sole cause of 2007-08 forecast failure.
- Sugihara et al. establish why correlation can mislead in nonlinear systems, but that paper is an ecological causality paper rather than a direct study of economic or epidemiological forecasting operations.
- The recommendation-system evidence shows drift-aware adaptation is useful, but it does not by itself quantify the exact share of degradation attributable to causal blindness versus interface, catalogue, or product changes.
Open Questions
- Which practical metrics best distinguish harmless recalibration from evidence that a model has lost contact with an invariant mechanism?
- How much explanatory structure is enough for operational robustness in settings where fully causal models are infeasible but pure prediction is brittle?
- What governance pattern is cheaper in practice: building more structural explanation into the model, or accepting instrumentalism and funding continual drift detection and repair?
sources
- [x] Friedman (1953) Essays in Positive Economics, Internet Archive record
- [x] Hausman (2018) Philosophy of Economics
- [x] Essays in Positive Economics - Wikipedia
- [x] Goodhart (1975) Problems of Monetary Management: The United Kingdom Experience, EconBiz record
- [x] Reserve Bank of Australia (1990) Conference volumes bibliography note on the origin of Goodhart's Law
- [x] Goodhart's law - Wikipedia
- [x] Pearl (2009) Causality excerpts
- [x] Pearl (2009) Causal inference in statistics: An overview
- [x] Pearl (2018) Theoretical Impediments to Machine Learning With Seven Sparks from the Causal Revolution
- [x] Shpitser and Pearl (2008) Complete Identification Methods for the Causal Hierarchy
- [x] Peters, Buhlmann, and Meinshausen (2015) Causal inference using invariant prediction: identification and confidence intervals
- [x] Buhlmann (2018) Invariance, Causality and Robustness
- [x] Arjovsky et al. (2019) Invariant Risk Minimization
- [x] Parascandolo et al. (2020) Learning explanations that are hard to vary
- [x] Sugihara et al. (2012) Detecting Causality in Complex Ecosystems
- [x] Casini and Perron (2018) Structural Breaks in Time Series
- [x] Rossi (2021) Forecasting in the Presence of Instabilities: How We Know Whether Models Predict Well and How to Improve Them
- [x] Bank for International Settlements (2008) Quarterly Review overview, December 2008
- [x] Shah et al. (2024) Accuracy of United States Centers for Disease Control and Prevention (US CDC) COVID-19 forecasting models
- [x] Nixon et al. (2022) Real-time COVID-19 forecasting: challenges and opportunities of model performance and translation
- [x] Hinder et al. (2024) One or two things we know about concept drift
- [x] Hartatik, Heryawan, and Pulungan (2025) Trust Decay-Based Temporal Learning for Dynamic Recommender Systems with Concept Drift Adaptation
- [x] Research Question 1.1 completed item
- [x] Research Question 1.2 completed item
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-19 | 4b9e57d | Initial completion |