−0.033R²
classical models forecast the future block better
Paired difference in one-step skill between the best neural and the best classical forecaster, each chosen on validation, across 8 sessions. 95% session-bootstrap interval [−0.053, −0.013]. Neural ahead in 1 of 8.
936
units analysed, visual cortex
8
sessions, one mouse each
0.187
best one-step R², AR (per unit)
+0.032
best free-rollout R² at 2.5 s, AR (per unit)
the question
how much do neural sequence models improve on classical dynamics when the test is the future, and a mouse they never saw
five hypotheses, frozen first
Each was fixed, with the rule that decides it, before the benchmark produced a result. The verdicts below are computed from the result files by that rule.
- H1
Neural sequence models beat linear AR on one-step forecasting.
not supported
effect −0.0328, 95% CI [−0.0529, −0.0127]
- H2
Classical models are competitive or better with little data.
supported
- H3
Neural-vs-linear gap narrows or reverses across sessions.
supported
effect −0.0040, 95% CI [−0.0052, −0.0022]
- H4
Long rollouts expose instability invisible at one step.
supported
- H5
Stimulus covariates improve prediction, mostly at long horizons, and mostly where there is a stimulus.
not supported
the evidence, in order
- 01dataset
Sessions, units, quality control and the split.
- 02activity
Recorded population activity beside each model's forecast.
- 03models
One-step benchmark with session-level uncertainty.
- 04forecast
Individual free rollouts, unit by unit.
- 05rollouts
Skill against horizon; teacher forcing against free rollout.
- 06latent
Dimension, stability and the aligned trajectories.
- 07generalization
Forecasting a mouse the model never saw.
- 08ablations
What the stimulus explains, and what it does not.
- 09failures
Where and how each model breaks.
- 10methods
Protocol, leakage controls and reproduction.
- 11paper
The manuscript, generated from the result files.