Formalising Popper's Falsifiability as a Mathematical Criterion for…
Formalising Popper's Falsifiability as a Mathematical Criterion for Distinguishing Mechanism from Interpolation
- Finite-sample Popperian falsifiability can be formalised as the amount of labeling space a hypothesis class excludes, because a class that realizes only `m_H(n)` of the `2^n` binary labelings on `n` points rules out the remaining possibilities and therefore makes riskier predictionsStanford (n.d.)Internet (n.d.)Wikipedia (n.d.)
- Vapnik-Chervonenkis dimension measures finite-sample expressive freedom rather than explanatory truth, so higher capacity weakens Popperian severity of test at a fixed sample size unless independent constraints shrink the realized growth functionStover (n.d.)Wikipedia (n.d.)Wikipedia (n.d.)
- Minimum Description Length operationalises Occam's Razor by selecting the hypothesis that minimises model code plus residual code, and Kolmogorov complexity supplies the ideal limiting notion of the shortest generative explanationRissanen (1978)Grunwald (2004)Vitanyi (2008)
- Mechanistic models differ from interpolators because they encode interpretable causal structure that travels beyond the fitted sample, whereas flexible machine-learning models can achieve strong prediction without exposing the underlying mechanismNielsen et al. (2025)Kording (2019)
- Modern deep networks show that interpolation and generalization can coexist, because overparameterized models can fit random labels or cross the interpolation threshold, the point at which training data can be fit exactly, and still recover lower test error afterwardZhang et al. (2017)Belkin et al. (2019)Nakkiran et al. (2020)
- A workable criterion between mechanism and interpolation is therefore conjunctive rather than binary-by-capacity: a model earns mechanistic status only when non-trivial logical content, short description length, and successful novelty testing all point in the same directionStanford (n.d.)Grunwald (2004)Nielsen et al. (2025)
- Under this criterion, unconstrained deep-learning models trained only for predictive accuracy should be treated as predictive tools rather than mechanism-level explanations unless architecture, symmetries, governing equations, or prior assumptions about causal structure sharply reduce effective freedom and explanation lengthBelkin et al. (2019)Nakkiran et al. (2020)Nielsen et al. (2025)Kording (2019)
Research Question
How can Karl Popper's criterion of demarcation and falsifiability be mathematically formalised to distinguish between a model that explains a physical mechanism and one that merely interpolates observational data?
Findings
(Populated from §6 Synthesis above.)
Executive Summary
A Popperian boundary between mechanistic explanation and interpolation can be formalised by combining finite-sample logical content, LC_n(H) = n - log2 m_H(n), with description length, DL(H,D) = L(H) + L(D|H), and under that combined test unconstrained deep networks do not earn mechanism-level status from interpolation or generalization alone.
Popper supplies the normative requirement that good theories forbid possibilities and survive severe tests, while Vapnik-Chervonenkis theory supplies a finite-sample count of remaining labelings and Minimum Description Length supplies a practical measure of explanation length.
The resulting criterion is conjunctive: a model is mechanistic only when it excludes many rival patterns, compresses the data with a short reusable description, and keeps working under novelty or intervention tests aimed at the claimed mechanism.
The main uncertainty is practical rather than conceptual: deep-learning-relevant capacity measures and code-length estimates are loose, so the framework is sharper as a decision rule for comparing model families than as a single scalar certificate for one trained network.
Key Findings
- Finite-sample Popperian falsifiability can be formalised as the amount of labeling space a hypothesis class excludes, because a class that realizes only
m_H(n)of the2^nbinary labelings onnpoints rules out the remaining possibilities and therefore makes riskier predictions. - Vapnik-Chervonenkis dimension measures finite-sample expressive freedom rather than explanatory truth, so higher capacity weakens Popperian severity of test at a fixed sample size unless independent constraints shrink the realized growth function.
- Minimum Description Length operationalises Occam's Razor by selecting the hypothesis that minimises model code plus residual code, and Kolmogorov complexity supplies the ideal limiting notion of the shortest generative explanation.
- Mechanistic models differ from interpolators because they encode interpretable causal structure that travels beyond the fitted sample, whereas flexible machine-learning models can achieve strong prediction without exposing the underlying mechanism.
- Modern deep networks show that interpolation and generalization can coexist, because overparameterized models can fit random labels or cross the interpolation threshold, the point at which training data can be fit exactly, and still recover lower test error afterward.
- A workable criterion between mechanism and interpolation is therefore conjunctive rather than binary-by-capacity: a model earns mechanistic status only when non-trivial logical content, short description length, and successful novelty testing all point in the same direction.
- Under this criterion, unconstrained deep-learning models trained only for predictive accuracy should be treated as predictive tools rather than mechanism-level explanations unless architecture, symmetries, governing equations, or prior assumptions about causal structure sharply reduce effective freedom and explanation length.
Assumptions
- [assumption] Exact Kolmogorov complexity is not required for the operational criterion, because practical model comparison can use code-length surrogates such as Minimum Description Length. [justification: the item asks for a usable mathematical formalisation rather than an incomputable ideal; source: Grunwald (2004) A tutorial introduction to the minimum description length principle doi.org
- [assumption] Finite-sample binary-labeling arguments are an acceptable proxy for Popper's excluded-observation logic even when the downstream physical problem uses real-valued observations. [justification: the learning-theory bridge requires a discrete count of possibilities before extending the intuition to richer observation spaces; source: Stanford Encyclopedia of Philosophy Karl Popper en.wikipedia.org
Analysis
The evidence supports a three-part construction rather than a single metric. Popper provides the normative intuition that a serious theory must exclude possibilities; Vapnik-Chervonenkis theory provides a finite-sample count of how many binary labelings a class still allows; Minimum Description Length provides a penalty for long fitted descriptions that merely memorize regularities.
A rival interpretation says that overparameterized deep networks can still discover real mechanisms because training dynamics may favor simpler solutions and architecture bias may recover low-dimensional structure. That rival remains plausible, but the accessible evidence here shows only that interpolation can coexist with generalization, not that the resulting representation is itself the underlying physical mechanism.
This is why raw Vapnik-Chervonenkis bounds are not the whole answer. Deep models often have huge classical capacity bounds, yet practical systems can still generalize. The correct conclusion is not that Popperian falsifiability fails, but that explanation requires extra evidence of compression and transport beyond the training regime.
The formal criterion is strongest when comparing rival model families on the same problem. If one family leaves many labelings open, needs a long code to describe itself, and fails on novelty tests, while another leaves fewer possibilities open, compresses the data with a shorter reusable code, and survives new tests, the second earns the stronger mechanistic claim.
Risks, Gaps, and Uncertainties
- The argument here relies on accessible secondary summaries for Popper and on accessible lecture-note or reference summaries for Vapnik-Chervonenkis theory, with the primary books kept as bibliographic locators rather than quoted chapter evidence.
- Classical Vapnik-Chervonenkis bounds are known to be loose for modern deep networks, so any criterion using them alone will overstate the case against neural models.
- Exact Kolmogorov complexity is not computable, so practical deployment of this framework must use proxy code lengths rather than the ideal shortest program.
- The framework is clearest for binary-labeled or explicitly coded prediction tasks and needs more work to handle continuous-time physical theories with rich intervention structure.
Open Questions
- How should effective dimension, margin bounds, or compression bounds replace raw Vapnik-Chervonenkis dimension when the model family is a modern transformer or diffusion architecture?
- What is the best practical coding scheme for measuring description length in fitted neural networks without smuggling in arbitrary engineering choices?
- Which intervention or regime-shift tests are decisive enough to let a statistical model earn a genuine mechanistic claim in physics or biology?
sources
- [ ] Popper (1959) The Logic of Scientific Discovery - primary source locator
- [ ] Vapnik (1995) The Nature of Statistical Learning Theory - primary source locator
- [x] Internet Encyclopedia of Philosophy Karl Popper: Philosophy of Science - accessible summary of falsifiability, demarcation, and ad hoc immunisation
- [x] Stanford Encyclopedia of Philosophy Karl Popper - accessible summary of falsifiability, empirical content, basic statements, and testability
- [x] Routledge The Logic of Scientific Discovery - chapter structure confirming falsifiability, degrees of testability, simplicity, and corroboration
- [x] Rissanen (1978) Modeling by shortest data description - abstract of the original Minimum Description Length proposal
- [x] Grunwald (2004) A tutorial introduction to the minimum description length principle - accessible technical tutorial on Minimum Description Length
- [x] Li and Vitanyi (2008) An Introduction to Kolmogorov Complexity and Its Applications - standard reference for Kolmogorov complexity and its relation to description length
- [x] Livni (2017) VC Dimension and the Fundamental Theorem - lecture note defining shattered sets and Vapnik-Chervonenkis dimension
- [x] Kakade (2020) Growth Functions and the VC dimension - lecture note defining the growth function and Sauer bound
- [x] Stover (n.d.) Vapnik-Chervonenkis Dimension - compact definition of shattering and Vapnik-Chervonenkis dimension
- [x] Wikipedia contributors (n.d.) Vapnik-Chervonenkis dimension - accessible definition of shattering and Vapnik-Chervonenkis dimension
- [x] Wikipedia contributors (n.d.) Vapnik-Chervonenkis theory - accessible description of growth functions and generalization context
- [x] Wikipedia contributors (n.d.) Minimum description length - accessible statement of the two-part code form
- [x] Wikipedia contributors (n.d.) Kolmogorov complexity - accessible definition of shortest-program complexity
- [x] Zhang et al. (2017) Understanding deep learning requires rethinking generalization - random-label fitting result
- [x] Belkin et al. (2019) Reconciling modern machine-learning practice and the classical bias-variance trade-off - double descent and interpolation
- [x] Nakkiran et al. (2020) Deep double descent: where bigger models and more data hurt - effective model complexity and empirical double descent
- [x] Nielsen et al. (2025) The rise of scientific machine learning - mechanistic models versus machine learning prediction
- [x] Lillicrap and Kording (2019) What does it mean to understand a neural network? - learning rules versus weight-level understanding
- [x] Research repo (2026-02-28) Predictive processing, active inference - prior repository item on falsifiability of a broad mathematical framework
- [x] Research repo (2026-02-28) Free energy, entropy, and life - prior repository item on framework-level versus process-level testability
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-19 | bb9188f | Initial completion |