Formalising Popper's Falsifiability as a Mathematical Criterion for…

Formalising Popper's Falsifiability as a Mathematical Criterion for Distinguishing Mechanism from Interpolation

2026-05-18 · ai-architecture benchmarks-eval epistemology-of-ai formal-learning-theory formal-methods · medium · source → · wiki →
key claims
  1. Finite-sample Popperian falsifiability can be formalised as the amount of labeling space a hypothesis class excludes, because a class that realizes only `m_H(n)` of the `2^n` binary labelings on `n` points rules out the remaining possibilities and therefore makes riskier predictionsStanford (n.d.)Internet (n.d.)Wikipedia (n.d.)
  2. Vapnik-Chervonenkis dimension measures finite-sample expressive freedom rather than explanatory truth, so higher capacity weakens Popperian severity of test at a fixed sample size unless independent constraints shrink the realized growth functionStover (n.d.)Wikipedia (n.d.)Wikipedia (n.d.)
  3. Minimum Description Length operationalises Occam's Razor by selecting the hypothesis that minimises model code plus residual code, and Kolmogorov complexity supplies the ideal limiting notion of the shortest generative explanationRissanen (1978)Grunwald (2004)Vitanyi (2008)
  4. Mechanistic models differ from interpolators because they encode interpretable causal structure that travels beyond the fitted sample, whereas flexible machine-learning models can achieve strong prediction without exposing the underlying mechanismNielsen et al. (2025)Kording (2019)
  5. Modern deep networks show that interpolation and generalization can coexist, because overparameterized models can fit random labels or cross the interpolation threshold, the point at which training data can be fit exactly, and still recover lower test error afterwardZhang et al. (2017)Belkin et al. (2019)Nakkiran et al. (2020)
  6. A workable criterion between mechanism and interpolation is therefore conjunctive rather than binary-by-capacity: a model earns mechanistic status only when non-trivial logical content, short description length, and successful novelty testing all point in the same directionStanford (n.d.)Grunwald (2004)Nielsen et al. (2025)
  7. Under this criterion, unconstrained deep-learning models trained only for predictive accuracy should be treated as predictive tools rather than mechanism-level explanations unless architecture, symmetries, governing equations, or prior assumptions about causal structure sharply reduce effective freedom and explanation lengthBelkin et al. (2019)Nakkiran et al. (2020)Nielsen et al. (2025)Kording (2019)

Research Question

How can Karl Popper's criterion of demarcation and falsifiability be mathematically formalised to distinguish between a model that explains a physical mechanism and one that merely interpolates observational data?

Findings

(Populated from §6 Synthesis above.)

Executive Summary

A Popperian boundary between mechanistic explanation and interpolation can be formalised by combining finite-sample logical content, LC_n(H) = n - log2 m_H(n), with description length, DL(H,D) = L(H) + L(D|H), and under that combined test unconstrained deep networks do not earn mechanism-level status from interpolation or generalization alone.

Popper supplies the normative requirement that good theories forbid possibilities and survive severe tests, while Vapnik-Chervonenkis theory supplies a finite-sample count of remaining labelings and Minimum Description Length supplies a practical measure of explanation length.

The resulting criterion is conjunctive: a model is mechanistic only when it excludes many rival patterns, compresses the data with a short reusable description, and keeps working under novelty or intervention tests aimed at the claimed mechanism.

The main uncertainty is practical rather than conceptual: deep-learning-relevant capacity measures and code-length estimates are loose, so the framework is sharper as a decision rule for comparing model families than as a single scalar certificate for one trained network.

Key Findings

  1. Finite-sample Popperian falsifiability can be formalised as the amount of labeling space a hypothesis class excludes, because a class that realizes only m_H(n) of the 2^n binary labelings on n points rules out the remaining possibilities and therefore makes riskier predictions.
  2. Vapnik-Chervonenkis dimension measures finite-sample expressive freedom rather than explanatory truth, so higher capacity weakens Popperian severity of test at a fixed sample size unless independent constraints shrink the realized growth function.
  3. Minimum Description Length operationalises Occam's Razor by selecting the hypothesis that minimises model code plus residual code, and Kolmogorov complexity supplies the ideal limiting notion of the shortest generative explanation.
  4. Mechanistic models differ from interpolators because they encode interpretable causal structure that travels beyond the fitted sample, whereas flexible machine-learning models can achieve strong prediction without exposing the underlying mechanism.
  5. Modern deep networks show that interpolation and generalization can coexist, because overparameterized models can fit random labels or cross the interpolation threshold, the point at which training data can be fit exactly, and still recover lower test error afterward.
  6. A workable criterion between mechanism and interpolation is therefore conjunctive rather than binary-by-capacity: a model earns mechanistic status only when non-trivial logical content, short description length, and successful novelty testing all point in the same direction.
  7. Under this criterion, unconstrained deep-learning models trained only for predictive accuracy should be treated as predictive tools rather than mechanism-level explanations unless architecture, symmetries, governing equations, or prior assumptions about causal structure sharply reduce effective freedom and explanation length.

Assumptions

Analysis

The evidence supports a three-part construction rather than a single metric. Popper provides the normative intuition that a serious theory must exclude possibilities; Vapnik-Chervonenkis theory provides a finite-sample count of how many binary labelings a class still allows; Minimum Description Length provides a penalty for long fitted descriptions that merely memorize regularities.

A rival interpretation says that overparameterized deep networks can still discover real mechanisms because training dynamics may favor simpler solutions and architecture bias may recover low-dimensional structure. That rival remains plausible, but the accessible evidence here shows only that interpolation can coexist with generalization, not that the resulting representation is itself the underlying physical mechanism.

This is why raw Vapnik-Chervonenkis bounds are not the whole answer. Deep models often have huge classical capacity bounds, yet practical systems can still generalize. The correct conclusion is not that Popperian falsifiability fails, but that explanation requires extra evidence of compression and transport beyond the training regime.

The formal criterion is strongest when comparing rival model families on the same problem. If one family leaves many labelings open, needs a long code to describe itself, and fails on novelty tests, while another leaves fewer possibilities open, compresses the data with a shorter reusable code, and survives new tests, the second earns the stronger mechanistic claim.

Risks, Gaps, and Uncertainties

Open Questions


sources


cites
cites Predictive processing and active inference: the brain as prediction machine
cites Free energy, entropy, and life: why organisms predict — from Schrödinger to Friston to Seth
related (frontmatter)
related Information synthesis: non-lossy compression, entropy, and information theory
related What are Barnum statements (Forer Effect statements), how do they manifest in Artificial Intelligence (AI)-generated text, and what methods exist to identify and remove them from AI research outputs?
version history
versiondatecommitsummary
1.02026-05-19bb9188fInitial completion

Connected items

Loading…

View full knowledge graph →