David Deutsch's Hard-to-Vary Criterion

David Deutsch's Hard-to-Vary Criterion: Measuring the Internal Logical Constraints of Explanatory Mechanisms

2026-05-18 · llm-reasoning ai-architecture benchmarks-eval epistemic-formalism model-interpretability formal-methods · medium · source → · wiki →
key claims
  1. Deutsch's hard-to-vary criterion complements the Popperian filter formalised in RQ 1.1 because it evaluates internal explanatory constraint before new evidence is gathered, whereas Popper evaluates excluded observations and risky tests after confrontation with dataResearch (2026)Stanford (n.d.)Parascandolo et al. (2020)
  2. Parascandolo et al. operationalise Deutsch's idea by defining consistency for minima across environments and by treating low-consistency minima as patchwork solutions that are likely to memorise rather than capture invariant mechanismsParascandolo et al. (2020)
  3. Variance-based sensitivity analysis supplies a measurable notion of internal logical coupling because the gap between a component's total effect and first-order effect quantifies how much explanatory work depends on coordinated interaction with other componentsSaltelli et al. (2010)Becker et al. (2017)
  4. Research on sloppiness, broad parameter directions that leave predictions nearly unchanged, and on structural identifiability, whether a model structure permits a unique parameter solution, shows that easy variation appears as broad parameter regions, whereas hard-to-vary explanations occupy comparatively small stiff regionsGutenkunst et al. (2007)Jagadeesan et al. (2023)Dufresne et al. (2018)
  5. A usable formal score for hard-to-vary-ness is therefore a composite of cross-environment consistency, interaction dominance, and admissible-variation volume, because no one of those quantities alone distinguishes genuine explanatory constraint from mere brittleness or mere good fitParascandolo et al. (2020)Saltelli et al. (2010)Jagadeesan et al. (2023)
  6. Varying internal constraints exposes structural fragility before empirical testing when the same fitted model can be rewritten through many compensating parameter changes, when minima disappear outside pooled data, or when component influence remains mostly separable rather than jointly constrainedParascandolo et al. (2020)Becker et al. (2017)Jagadeesan et al. (2023)
  7. In the contrast reviewed here, flexible deep neural networks trained only for prediction are more likely than compact mechanistic models to score lower on this criterion, while hybrid scientific machine learning systems can raise their score when architecture and training explicitly enforce governing structure or cross-environment invarianceNielsen et al. (2025)Parascandolo et al. (2020)

Research Question

Using David Deutsch's hard-to-vary criterion, meaning an explanation whose details cannot be changed without losing explanatory force, what formal criteria can measure the internal logical constraints of an explanatory mechanism, and how does varying those constraints expose structural fragility before empirical testing occurs?

Findings

(Populated from §6 Synthesis above.)

Executive Summary

Deutsch's hard-to-vary criterion can be formalised as a joint test of cross-environment consistency, interaction-dominated sensitivity, and small explanation-preserving parameter volume, rather than as any single metric taken in isolation.

An explanation is hard to vary when the same minimum or mechanism recurs across environments, when most component influence is carried by interactions with other components, and when only a small region of plausible parameter space preserves explanatory coherence.

Varying those internal constraints exposes structural fragility before new empirical testing because low consistency reveals patchwork solutions, low interaction dominance reveals independently tunable parts, and large admissible variation regions reveal many compensating rewrites that leave current fit intact.

This criterion complements rather than replaces the Popperian filter in RQ 1.1: Popper still decides the post-empirical standing of a theory, while the hard-to-vary score decides whether the explanation is internally constrained enough to justify serious empirical investment.

Key Findings

  1. Deutsch's hard-to-vary criterion complements the Popperian filter formalised in RQ 1.1 because it evaluates internal explanatory constraint before new evidence is gathered, whereas Popper evaluates excluded observations and risky tests after confrontation with data.
  2. Parascandolo et al. operationalise Deutsch's idea by defining consistency for minima across environments and by treating low-consistency minima as patchwork solutions that are likely to memorise rather than capture invariant mechanisms.
  3. Variance-based sensitivity analysis supplies a measurable notion of internal logical coupling because the gap between a component's total effect and first-order effect quantifies how much explanatory work depends on coordinated interaction with other components.
  4. Research on sloppiness, broad parameter directions that leave predictions nearly unchanged, and on structural identifiability, whether a model structure permits a unique parameter solution, shows that easy variation appears as broad parameter regions, whereas hard-to-vary explanations occupy comparatively small stiff regions.
  5. A usable formal score for hard-to-vary-ness is therefore a composite of cross-environment consistency, interaction dominance, and admissible-variation volume, because no one of those quantities alone distinguishes genuine explanatory constraint from mere brittleness or mere good fit.
  6. Varying internal constraints exposes structural fragility before empirical testing when the same fitted model can be rewritten through many compensating parameter changes, when minima disappear outside pooled data, or when component influence remains mostly separable rather than jointly constrained.
  7. In the contrast reviewed here, flexible deep neural networks trained only for prediction are more likely than compact mechanistic models to score lower on this criterion, while hybrid scientific machine learning systems can raise their score when architecture and training explicitly enforce governing structure or cross-environment invariance.

Assumptions

Analysis

The evidence supports a three-part criterion rather than a single statistic, because cross-environment recurrence, interaction-mediated dependence, and small admissible variation regions each capture a different way in which an explanation resists arbitrary rewriting.

A rival interpretation says that any model with high local sensitivity is already hard to vary, but that interpretation confuses brittleness with explanatory constraint because a one-parameter unstable fit can still be easy to rewrite elsewhere in parameter space.

Another rival interpretation says that mechanistic models are always hard to vary and neural models are always easy to vary, but the scientific machine learning literature shows that hybrid models can acquire genuine internal constraint when sparse architecture, governing equations, or prior structure limit arbitrary rewrites.

The proposed HTV score should therefore be used comparatively across rival model families or rival training setups, not as a metaphysical certificate that one fitted parameter vector has captured reality once and for all.

Risks, Gaps, and Uncertainties

Open Questions


sources


cites
cites Formalising Popper's Falsifiability as a Mathematical Criterion for Distinguishing Mechanism from Interpolation
related (frontmatter)
related Formalising Popper's Falsifiability as a Mathematical Criterion for Distinguishing Mechanism from Interpolation
related Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic Systems Under Distribution Shift
version history
versiondatecommitsummary
1.02026-05-19f0637b1Initial completion

Connected items

Loading…

View full knowledge graph →