David Deutsch's Hard-to-Vary Criterion
David Deutsch's Hard-to-Vary Criterion: Measuring the Internal Logical Constraints of Explanatory Mechanisms
- Deutsch's hard-to-vary criterion complements the Popperian filter formalised in RQ 1.1 because it evaluates internal explanatory constraint before new evidence is gathered, whereas Popper evaluates excluded observations and risky tests after confrontation with dataResearch (2026)Stanford (n.d.)Parascandolo et al. (2020)
- Parascandolo et al. operationalise Deutsch's idea by defining consistency for minima across environments and by treating low-consistency minima as patchwork solutions that are likely to memorise rather than capture invariant mechanismsParascandolo et al. (2020)
- Variance-based sensitivity analysis supplies a measurable notion of internal logical coupling because the gap between a component's total effect and first-order effect quantifies how much explanatory work depends on coordinated interaction with other componentsSaltelli et al. (2010)Becker et al. (2017)
- Research on sloppiness, broad parameter directions that leave predictions nearly unchanged, and on structural identifiability, whether a model structure permits a unique parameter solution, shows that easy variation appears as broad parameter regions, whereas hard-to-vary explanations occupy comparatively small stiff regionsGutenkunst et al. (2007)Jagadeesan et al. (2023)Dufresne et al. (2018)
- A usable formal score for hard-to-vary-ness is therefore a composite of cross-environment consistency, interaction dominance, and admissible-variation volume, because no one of those quantities alone distinguishes genuine explanatory constraint from mere brittleness or mere good fitParascandolo et al. (2020)Saltelli et al. (2010)Jagadeesan et al. (2023)
- Varying internal constraints exposes structural fragility before empirical testing when the same fitted model can be rewritten through many compensating parameter changes, when minima disappear outside pooled data, or when component influence remains mostly separable rather than jointly constrainedParascandolo et al. (2020)Becker et al. (2017)Jagadeesan et al. (2023)
- In the contrast reviewed here, flexible deep neural networks trained only for prediction are more likely than compact mechanistic models to score lower on this criterion, while hybrid scientific machine learning systems can raise their score when architecture and training explicitly enforce governing structure or cross-environment invarianceNielsen et al. (2025)Parascandolo et al. (2020)
Research Question
Using David Deutsch's hard-to-vary criterion, meaning an explanation whose details cannot be changed without losing explanatory force, what formal criteria can measure the internal logical constraints of an explanatory mechanism, and how does varying those constraints expose structural fragility before empirical testing occurs?
Findings
(Populated from §6 Synthesis above.)
Executive Summary
Deutsch's hard-to-vary criterion can be formalised as a joint test of cross-environment consistency, interaction-dominated sensitivity, and small explanation-preserving parameter volume, rather than as any single metric taken in isolation.
An explanation is hard to vary when the same minimum or mechanism recurs across environments, when most component influence is carried by interactions with other components, and when only a small region of plausible parameter space preserves explanatory coherence.
Varying those internal constraints exposes structural fragility before new empirical testing because low consistency reveals patchwork solutions, low interaction dominance reveals independently tunable parts, and large admissible variation regions reveal many compensating rewrites that leave current fit intact.
This criterion complements rather than replaces the Popperian filter in RQ 1.1: Popper still decides the post-empirical standing of a theory, while the hard-to-vary score decides whether the explanation is internally constrained enough to justify serious empirical investment.
Key Findings
- Deutsch's hard-to-vary criterion complements the Popperian filter formalised in RQ 1.1 because it evaluates internal explanatory constraint before new evidence is gathered, whereas Popper evaluates excluded observations and risky tests after confrontation with data.
- Parascandolo et al. operationalise Deutsch's idea by defining consistency for minima across environments and by treating low-consistency minima as patchwork solutions that are likely to memorise rather than capture invariant mechanisms.
- Variance-based sensitivity analysis supplies a measurable notion of internal logical coupling because the gap between a component's total effect and first-order effect quantifies how much explanatory work depends on coordinated interaction with other components.
- Research on sloppiness, broad parameter directions that leave predictions nearly unchanged, and on structural identifiability, whether a model structure permits a unique parameter solution, shows that easy variation appears as broad parameter regions, whereas hard-to-vary explanations occupy comparatively small stiff regions.
- A usable formal score for hard-to-vary-ness is therefore a composite of cross-environment consistency, interaction dominance, and admissible-variation volume, because no one of those quantities alone distinguishes genuine explanatory constraint from mere brittleness or mere good fit.
- Varying internal constraints exposes structural fragility before empirical testing when the same fitted model can be rewritten through many compensating parameter changes, when minima disappear outside pooled data, or when component influence remains mostly separable rather than jointly constrained.
- In the contrast reviewed here, flexible deep neural networks trained only for prediction are more likely than compact mechanistic models to score lower on this criterion, while hybrid scientific machine learning systems can raise their score when architecture and training explicitly enforce governing structure or cross-environment invariance.
Assumptions
- [assumption] The consistency term
C(M)can be approximated by environment splits or other meaningful partitions of data rather than by access to every possible deployment environment. [justification: Parascandolo's operationalisation already uses multiple environments as the relevant testbed for invariance; source: arxiv.org/abs/2009.00329] - [assumption] The admissible-volume term
V(M)can be estimated with local Fisher or Hessian geometry and identifiability diagnostics even when the exact global volume is computationally impractical to calculate. [justification: the sloppiness literature treats induced local geometry as a practical route to diagnosing broad versus narrow parameter directions; source: Dufresne et al. (2018) The geometry of sloppiness pmc.ncbi.nlm.nih.gov - [assumption] Interaction dominance from variance-based sensitivity indices is an acceptable proxy for internal logical coupling, even though logical dependence in the philosophical sense is richer than variance decomposition alone. [justification: the sensitivity literature measures whether factor influence is isolated or interaction-mediated, which is the operational distinction needed for this item; source: Saltelli et al. (2010) Variance based sensitivity analysis of model output. Design and estimator for the total sensitivity index pmc.ncbi.nlm.nih.gov
Analysis
The evidence supports a three-part criterion rather than a single statistic, because cross-environment recurrence, interaction-mediated dependence, and small admissible variation regions each capture a different way in which an explanation resists arbitrary rewriting.
A rival interpretation says that any model with high local sensitivity is already hard to vary, but that interpretation confuses brittleness with explanatory constraint because a one-parameter unstable fit can still be easy to rewrite elsewhere in parameter space.
Another rival interpretation says that mechanistic models are always hard to vary and neural models are always easy to vary, but the scientific machine learning literature shows that hybrid models can acquire genuine internal constraint when sparse architecture, governing equations, or prior structure limit arbitrary rewrites.
The proposed HTV score should therefore be used comparatively across rival model families or rival training setups, not as a metaphysical certificate that one fitted parameter vector has captured reality once and for all.
Risks, Gaps, and Uncertainties
- The consulted public source for Deutsch's book is a publisher page that confirms the work and its focus on explanations, but it does not expose the chapter text, so phrase-level interpretation depends on Parascandolo's accessible quotation of the criterion.
- Variance-based sensitivity metrics depend on the chosen output variable and on the assumed input-distribution or uncertainty structure, so the same mechanism can receive different interaction scores under different formal problem statements.
- Local sloppiness geometry can miss non-local compensating variations, so practical estimation of admissible volume is an approximation rather than a complete global certificate.
- The proposed score is sharper for comparing rival explanatory models on the same task than for assigning a universally meaningful absolute threshold that separates explanation from non-explanation across all domains.
Open Questions
- What is the best practical estimator for admissible explanation-preserving volume in large neural systems where the local Hessian badly under-represents the global solution manifold?
- How should the interaction term be adapted for explanations whose key constraints are structural or symbolic rather than parametrically differentiable?
- Which intervention or regime-shift tests are stringent enough to convert a high pre-empirical
HTVscore into the stronger post-empirical status formalised in RQ 1.1?
sources
- [x] Deutsch (2011) The Beginning of Infinity - publisher page used as the primary locator for the book that introduced the hard-to-vary criterion.
- [x] Parascandolo et al. (2020) Learning Explanations That Are Hard to Vary - direct mathematical operationalisation of the criterion through consistency across environments.
- [x] Saltelli et al. (2008) Global Sensitivity Analysis: The Primer - official Wiley landing page for the seeded sensitivity-analysis book.
- [x] Saltelli et al. (2010) Variance based sensitivity analysis of model output. Design and estimator for the total sensitivity index - primary summary of first-order and total-effect variance decomposition.
- [x] Becker et al. (2017) Exploring the weighting effects of composite indicators - accessible treatment of the correlation ratio as a first-order sensitivity index and of conditional variance reduction.
- [x] Gutenkunst et al. (2007) Universally Sloppy Parameter Sensitivities in Systems Biology Models - classic statement of sloppy and stiff parameter directions.
- [x] Jagadeesan et al. (2023) Sloppiness: fundamental study, new formalism and quantification - explicit connection among sloppiness, structural identifiability, and nearly identical prediction regions.
- [x] Dufresne et al. (2018) The geometry of sloppiness - formal treatment of sloppiness via an induced premetric on parameter space.
- [x] Nielsen et al. (2025) The rise of scientific machine learning - current review on mechanistic models, machine learning, and hybrid scientific machine learning systems.
- [x] Research repo (2026) Research Question 1.1 (RQ 1.1): Formalising Popper's Falsifiability as a Criterion Between Mechanism and Interpolation - directly preceding item that formalised the post-empirical filter this item extends.
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-19 | f0637b1 | Initial completion |