15 public preprints. Every one is pre-registered with kill conditions committed before results, and every reported number regenerates from committed run artifacts. DOI links resolve to the latest version on Zenodo.
What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
Characterizes, across 188 controlled runs, when a representational prior accelerates grokking: the prior must be built from the right feature family, label-free invariances outperform supervised signals, and the benefit is concentrated in a narrow critical window early in training.
Structure-Specific Representational Priors Causally Control the Grokking Delay
Controlled experiments showing the grokking delay is causally the time to form task-structured representations. Injecting the right structure through a contrastive representational prior collapses the delay — up to 22× faster generalization — while structure-agnostic interventions do not.
Contracting, Non-Normal, and Not the Mechanism: A Pre-Registered Spectral Audit of Latent Chain-of-Thought Reasoning
Treats a Coconut-style latent reasoner's chain of continuous thoughts as a discrete-time dynamical system and pre-registers spectral predictions with kill conditions. The kill condition fired: the dynamics are contracting and non-normal, but spectral invariants are not the mechanism of latent reasoning.
Same Output, Different Process: Ornstein's d̄-Distance as a Computational-Equivalence Metric for Model Diffing
When are two neural networks the same computation? A process-level equivalence metric from ergodic theory that separates what representation-geometry tools (CKA, SVCCA) and linearized-dynamics tools (DSA) conflate — such as a model versus its pruned, distilled, or noise-injected twin.
Intrinsic-Noise Consolidation: A Doob-Barrier-Conditioned Diffusion Turns Analog Device Noise into a Continual-Learning Resource
Casts per-synapse memory consolidation as a Doob h-transform: condition each weight's stochastic dynamics on never crossing a memory barrier, and intrinsic device noise becomes a consolidation resource instead of an accuracy tax. Confirmed on real BrainScaleS-2 analog neuromorphic hardware, with +15.6 points retention.
Heckman-Corrected Epistemic Uncertainty: Selection on Unobservables Defeats Importance Weighting
Training data is routinely collected by a selection process the model never sees — loans observed only when granted, outcomes only when a test was ordered. When selection acts on unobservables, importance weighting and covariate-shift corrections fail; adapting Heckman's econometric selection model to deep networks restores calibrated uncertainty.
Survivor Bias in Learning-Curve Surrogates: Successive Halving Selects on Noise
Learning-curve surrogates for hyperparameter optimization are fit on curves that survived earlier rungs of successive halving — a censoring mechanism no method in that line models. The paper quantifies the resulting survivor bias and treats it with selection-aware corrections.
Interpolation Beats Finite-Size Scaling for Domain Extension of Learned Transfer-Operator Spectra: A Two-System Study and a Convergence Diagnostic
Learns Koopman/transfer-operator spectra on small domains of chaotic PDEs and extrapolates to domains up to 64× larger. The finite-size-scaling pipeline works in absolute terms — but never beats simple interpolation. An honest negative, shipped with a practical convergence diagnostic.
Rice's Formula for Event-Driven Networks: Predicting and Training Temporal Sparsity via Differentiable Level-Crossing Rates
A delta-network event is, definitionally, a level crossing of an activation time series — so Rice's formula predicts event rates from activation statistics. This yields closed-form threshold allocation and a differentiable handle on temporal sparsity, with a pre-registered negative on end-to-end budget training reported alongside.
A One-Sided Level-Crossing Budget for Physics-Informed Neural Networks: Validated Mechanism, Absent Pathology
Builds a one-sided Kac–Rice level-crossing budget for PINNs on stiff PDEs and validates its mechanism — then reports that the widely-cited catastrophic high-frequency divergence it was designed to prevent does not appear in a careful reproduction of the benchmark.
Differentiable Minkowski Functionals for Neural Fields: Validation, Design Rules, and an Adversarial Failure Mode
Smooth Monte-Carlo estimators for the complete integral-geometric description of a neural field's level sets — area, boundary measure, and Euler characteristic — enabling topology repair roughly 250× faster than persistent-homology pipelines, with validated design rules and a documented adversarial failure mode.
Event-Timing Priors: Temporal Point Process Likelihoods Correct the Long-Horizon Extreme-Event Statistics of Generative Surrogates
Generative surrogates of chaotic systems — including diffusion-based weather and climate emulators — are evaluated on extreme events but never trained on their timing. Adding a temporal-point-process likelihood over event times corrects long-horizon extreme-event statistics on a two-scale Lorenz-96 testbed.
Level-Crossing Density as a Mesh-Free High-Frequency Auxiliary Loss for Implicit Neural Representations
Coordinate-MLP neural fields exhibit spectral bias: low frequencies fit quickly, high frequencies slowly or never. The Kac–Rice level-crossing density gives a mesh-free, FFT-free auxiliary loss that acts directly on scattered data in the spatial domain — where frequency-domain remedies require grids.
Ornstein's d̄-Distance as a Dynamical-Fidelity Metric for Surrogate Models of Chaotic Systems
Two chaotic systems can share an invariant measure yet compute different dynamics — a failure Wasserstein-style evaluation cannot see. Ornstein's d̄-distance from ergodic isomorphism theory catches it, validated on Lorenz-63 and Kuramoto–Sivashinsky surrogates.
Semantic Reference Frames: Representing Time and Context Relative to Learned Anchors in Sequence Models
Positional signals tied to absolute indices saturate over long spans and generalize poorly beyond training length. Semantic reference frames measure position relative to learned semantic anchors — a frame-conditioned attention bias that generalizes ALiBi and improves length extrapolation on recall tasks and enwik8.