The Duhem-Quine Thesis and Underdetermination

The Duhem-Quine Thesis and Underdetermination: Quantifying When a Model Has Matched the True Mechanism vs. an Observational Proxy

2026-05-18 · ai-architecture benchmarks-eval epistemic-foundations causal-modeling · medium · source → · wiki →
key claims
  1. The Duhem-Quine thesis maps directly onto model selection because finite evidence constrains a web of hypotheses rather than selecting a unique mechanism, so several rival theories can remain compatible with exactly the same observationsQuine (1951)Philosophy (2017)
  2. Polynomial interpolation is unique only inside a restricted hypothesis class, degree at most `n` polynomials for `n+1` nodes, which shows that uniqueness comes from model-class assumptions rather than from the observations aloneJensen (2024)
  3. Bongard and Lipson's active-probing framework shows that candidate dynamical models can fit the same current observations and only become distinguishable after new perturbations or broader observability reveal divergent predictionsLipson (2007)
  4. Raue et al. provide a quantitative test by defining structural non-identifiability as unchanged observables under redundant parameterisations and practical non-identifiability as unbounded likelihood-based confidence regions caused by insufficient dataRaue et al. (2009)
  5. A model has matched the true mechanism only if mechanism-bearing parameters are structurally identifiable and practically bounded, and if interventions or environment changes fail to expose a rival predictor with the same observational fitRaue et al. (2009)Meinshausen (2015)Pearl (2018)
  6. Causal-invariance results explain why low observational risk is insufficient, because predictors built on non-causal associations can match training data while breaking under interventions or distribution shiftMeinshausen (2015)Pearl (2018)Scholkopf et al. (2021)
  7. Standard deep neural networks do not by themselves guarantee mechanism recovery because multiple parameterisations can implement the same function and the available identifiability results are explicitly architecture-specific and symmetry-boundedBona-Pellissier et al. (2023)Shen (2024)
  8. Deep learning can approach mechanism identification only when passive fitting is supplemented by restrictive architecture, sufficient excitation of relevant inputs, and multi-environment or interventional evidence that breaks proxy equivalence classesBona-Pellissier et al. (2023)Meinshausen (2015)Scholkopf et al. (2021)Lipson (2007)

Research Question

How can the phenomenon of multiple distinct functions perfectly interpolating identical data points be formalised through the lens of the Duhem-Quine thesis, underdetermination of theory by data, and what are the quantitative metrics for evaluating when a model has matched the true mechanism rather than an observationally equivalent proxy?

Findings

Executive Summary

Finite observational data do not justify the claim that a model has matched the true mechanism; they justify at most that the model belongs to an equivalence class of theories consistent with the observed traces unless identifiability and invariance conditions are also satisfied.

In dynamical systems, structural identifiability asks whether distinct parameter settings can leave observables unchanged, while practical identifiability asks whether finite noisy data leave profile-likelihood confidence regions unbounded.

A model counts as mechanism matched only when mechanism-bearing parameters are structurally identifiable, practically constrained, and supported by intervention or multi-environment evidence that rules out observationally equivalent proxies.

Standard deep-learning models trained on passive data do not establish mechanism matching by default because functional and parameter symmetries leave equivalent fits alive and observational risk alone does not certify causal structure.

Key Findings

  1. The Duhem-Quine thesis maps directly onto model selection because finite evidence constrains a web of hypotheses rather than selecting a unique mechanism, so several rival theories can remain compatible with exactly the same observations.
  2. Polynomial interpolation is unique only inside a restricted hypothesis class, degree at most n polynomials for n+1 nodes, which shows that uniqueness comes from model-class assumptions rather than from the observations alone.
  3. Bongard and Lipson's active-probing framework shows that candidate dynamical models can fit the same current observations and only become distinguishable after new perturbations or broader observability reveal divergent predictions.
  4. Raue et al. provide a quantitative test by defining structural non-identifiability as unchanged observables under redundant parameterisations and practical non-identifiability as unbounded likelihood-based confidence regions caused by insufficient data.
  5. A model has matched the true mechanism only if mechanism-bearing parameters are structurally identifiable and practically bounded, and if interventions or environment changes fail to expose a rival predictor with the same observational fit.
  6. Causal-invariance results explain why low observational risk is insufficient, because predictors built on non-causal associations can match training data while breaking under interventions or distribution shift.
  7. Standard deep neural networks do not by themselves guarantee mechanism recovery because multiple parameterisations can implement the same function and the available identifiability results are explicitly architecture-specific and symmetry-bounded.
  8. Deep learning can approach mechanism identification only when passive fitting is supplemented by restrictive architecture, sufficient excitation of relevant inputs, and multi-environment or interventional evidence that breaks proxy equivalence classes.

Assumptions

Analysis

The evidence weighs most strongly on the negative claim, observational fit alone is insufficient, because Quine, Raue, Pearl, Peters et al., and Scholkopf et al. all converge on the same asymmetry between observed agreement and mechanism-level warrant.

The positive criterion must therefore be conjunctive rather than singular: identifiability without intervention sensitivity can still miss proxy mechanisms, while invariance claims without identifiability can still hide redundant parameterisations.

Bongard and Lipson sharpen this by showing that experiment design, not just loss minimisation, determines whether rival mechanisms remain observationally equivalent, which makes active probing part of the epistemic test rather than a downstream optimisation detail.

The deep-learning case remains more conditional than the dynamical-systems case because the identifiable special results are architecture-specific and symmetry-bounded, whereas large practical training pipelines rarely satisfy those restrictive premises transparently.

Risks, Gaps, and Uncertainties

Open Questions


sources


cites
cites Empirical Risk Minimisation's Causal Blindness: Why In-Distribution Accuracy Guarantees Break Under Environment Change
cites Formalising Popper's Falsifiability as a Mathematical Criterion for Distinguishing Mechanism from Interpolation
cites Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic Systems Under Distribution Shift
related (frontmatter)
related Structural Stability vs. Predictive Fragility: Dynamical Systems Theory and the Cost of Noise in Mechanism-Free Models
related Pearl's Causal Hierarchy: Formal Information-Theoretic Limits on Deriving Interventional and Counterfactual Reasoning from Observational Data
version history
versiondatecommitsummary
1.02026-05-19e1e4be7Initial completion

Connected items

Loading…

View full knowledge graph →