The Duhem-Quine Thesis and Underdetermination
The Duhem-Quine Thesis and Underdetermination: Quantifying When a Model Has Matched the True Mechanism vs. an Observational Proxy
- The Duhem-Quine thesis maps directly onto model selection because finite evidence constrains a web of hypotheses rather than selecting a unique mechanism, so several rival theories can remain compatible with exactly the same observationsQuine (1951)Philosophy (2017)
- Polynomial interpolation is unique only inside a restricted hypothesis class, degree at most `n` polynomials for `n+1` nodes, which shows that uniqueness comes from model-class assumptions rather than from the observations aloneJensen (2024)
- Bongard and Lipson's active-probing framework shows that candidate dynamical models can fit the same current observations and only become distinguishable after new perturbations or broader observability reveal divergent predictionsLipson (2007)
- Raue et al. provide a quantitative test by defining structural non-identifiability as unchanged observables under redundant parameterisations and practical non-identifiability as unbounded likelihood-based confidence regions caused by insufficient dataRaue et al. (2009)
- A model has matched the true mechanism only if mechanism-bearing parameters are structurally identifiable and practically bounded, and if interventions or environment changes fail to expose a rival predictor with the same observational fitRaue et al. (2009)Meinshausen (2015)Pearl (2018)
- Causal-invariance results explain why low observational risk is insufficient, because predictors built on non-causal associations can match training data while breaking under interventions or distribution shiftMeinshausen (2015)Pearl (2018)Scholkopf et al. (2021)
- Standard deep neural networks do not by themselves guarantee mechanism recovery because multiple parameterisations can implement the same function and the available identifiability results are explicitly architecture-specific and symmetry-boundedBona-Pellissier et al. (2023)Shen (2024)
- Deep learning can approach mechanism identification only when passive fitting is supplemented by restrictive architecture, sufficient excitation of relevant inputs, and multi-environment or interventional evidence that breaks proxy equivalence classesBona-Pellissier et al. (2023)Meinshausen (2015)Scholkopf et al. (2021)Lipson (2007)
Research Question
How can the phenomenon of multiple distinct functions perfectly interpolating identical data points be formalised through the lens of the Duhem-Quine thesis, underdetermination of theory by data, and what are the quantitative metrics for evaluating when a model has matched the true mechanism rather than an observationally equivalent proxy?
Findings
Executive Summary
Finite observational data do not justify the claim that a model has matched the true mechanism; they justify at most that the model belongs to an equivalence class of theories consistent with the observed traces unless identifiability and invariance conditions are also satisfied.
In dynamical systems, structural identifiability asks whether distinct parameter settings can leave observables unchanged, while practical identifiability asks whether finite noisy data leave profile-likelihood confidence regions unbounded.
A model counts as mechanism matched only when mechanism-bearing parameters are structurally identifiable, practically constrained, and supported by intervention or multi-environment evidence that rules out observationally equivalent proxies.
Standard deep-learning models trained on passive data do not establish mechanism matching by default because functional and parameter symmetries leave equivalent fits alive and observational risk alone does not certify causal structure.
Key Findings
- The Duhem-Quine thesis maps directly onto model selection because finite evidence constrains a web of hypotheses rather than selecting a unique mechanism, so several rival theories can remain compatible with exactly the same observations.
- Polynomial interpolation is unique only inside a restricted hypothesis class, degree at most
npolynomials forn+1nodes, which shows that uniqueness comes from model-class assumptions rather than from the observations alone. - Bongard and Lipson's active-probing framework shows that candidate dynamical models can fit the same current observations and only become distinguishable after new perturbations or broader observability reveal divergent predictions.
- Raue et al. provide a quantitative test by defining structural non-identifiability as unchanged observables under redundant parameterisations and practical non-identifiability as unbounded likelihood-based confidence regions caused by insufficient data.
- A model has matched the true mechanism only if mechanism-bearing parameters are structurally identifiable and practically bounded, and if interventions or environment changes fail to expose a rival predictor with the same observational fit.
- Causal-invariance results explain why low observational risk is insufficient, because predictors built on non-causal associations can match training data while breaking under interventions or distribution shift.
- Standard deep neural networks do not by themselves guarantee mechanism recovery because multiple parameterisations can implement the same function and the available identifiability results are explicitly architecture-specific and symmetry-bounded.
- Deep learning can approach mechanism identification only when passive fitting is supplemented by restrictive architecture, sufficient excitation of relevant inputs, and multi-environment or interventional evidence that breaks proxy equivalence classes.
Assumptions
- "Mechanism-bearing parameters" means the subset of parameters whose variation changes intervention-relevant or environment-invariant behaviour, not merely a redundant reparameterisation of the same observational map.
- The target mechanism is representable within the candidate model class being evaluated, because identifiability results cannot recover a mechanism that the class cannot express.
Analysis
The evidence weighs most strongly on the negative claim, observational fit alone is insufficient, because Quine, Raue, Pearl, Peters et al., and Scholkopf et al. all converge on the same asymmetry between observed agreement and mechanism-level warrant.
The positive criterion must therefore be conjunctive rather than singular: identifiability without intervention sensitivity can still miss proxy mechanisms, while invariance claims without identifiability can still hide redundant parameterisations.
Bongard and Lipson sharpen this by showing that experiment design, not just loss minimisation, determines whether rival mechanisms remain observationally equivalent, which makes active probing part of the epistemic test rather than a downstream optimisation detail.
The deep-learning case remains more conditional than the dynamical-systems case because the identifiable special results are architecture-specific and symmetry-bounded, whereas large practical training pipelines rarely satisfy those restrictive premises transparently.
Risks, Gaps, and Uncertainties
- Structural identifiability is model-class relative, so a model can be identifiable inside a misspecified class and still fail to capture the real-world mechanism outside that class.
- The deep-learning conclusion is medium confidence rather than high confidence because the cited identifiability results cover restricted network families, not the full range of contemporary large-scale training settings.
- Intervention-based criteria are strongest when feasible, but many practical domains still rely on partial observability or limited perturbation budgets, which leaves some equivalence classes unresolved even after careful measurement design.
Open Questions
- Which practical benchmark family best measures mechanism identification, rather than interpolation, in modern deep-learning systems?
- How should identifiability be defined when the target is a latent causal representation rather than an explicit ordinary differential equation or graph parameter set?
- What minimum intervention or environment-variation budget is sufficient to break the most important proxy equivalence classes in foundation-model settings?
sources
- [x] Quine (1951) Two Dogmas of Empiricism, accessible text - primary source for Quine's confirmation holism, especially the claim that statements face experience as a corporate body.
- [x] Stanford Encyclopedia of Philosophy (2017) Scientific Underdetermination - authoritative secondary source distinguishing holist and contrastive underdetermination and locating Duhem and Quine within that debate.
- [x] Bongard and Lipson (2007) Automated reverse engineering of nonlinear dynamical systems - accessible full text for active probing, candidate-model disagreement, and symbolic mechanism discovery in nonlinear ordinary differential equation (ODE) systems.
- [x] Raue et al. (2009) Structural and practical identifiability analysis of partially observed dynamical models by exploiting the profile likelihood - accessible full text for structural identifiability, practical identifiability, and finite confidence-interval criteria.
- [x] Peters, Buhlmann, and Meinshausen (2015) Causal inference using invariant prediction: identification and confidence intervals - accessible source for invariance as a mechanism-sensitive criterion under interventions and environment change.
- [x] Pearl (2018) Theoretical Impediments to Machine Learning With Seven Sparks from the Causal Revolution - accessible source on the limit of purely associational learning for intervention questions.
- [x] Scholkopf et al. (2021) Toward Causal Representation Learning - accessible review connecting independently and identically distributed (i.i.d.) learning, interventions, and causal mechanisms.
- [x] Bona-Pellissier et al. (2023) Parameter identifiability of a deep feedforward rectified linear unit (ReLU) neural network - accessible source on when deep-network parameters are identifiable modulo known symmetries.
- [x] Shen (2024) Exploring the Complexity of Deep Neural Networks through Functional Equivalence - accessible source showing that different parameterisations can implement the same neural-network function.
- [x] Jensen (2024) Computational Methods lecture notes, polynomial interpolation - accessible source for the existence and uniqueness theorem inside the restricted polynomial class.
- [x] Research repo (2026-05-18) Research Question 2.1: Empirical Risk Minimisation's causal blindness and the limits of in-distribution guarantees - directly preceding item that this question sharpens.
- [x] Research repo (2026-05-18) Research Question 1.1: Formalising Popper's falsifiability as a criterion between mechanism and interpolation - prior repository item on mechanism versus interpolation.
- [x] Research repo (2026-05-18) Research Question 1.3: Failure modes of instrumentalism when applied to complex dynamic systems under distribution shift - prior repository item on predictive success without mechanistic robustness.
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-19 | e1e4be7 | Initial completion |