where it breaks
Eight ways a forecaster can fail while its average error still looks fine. Each has a statistic and a threshold fixed before the final runs, so the taxonomy cannot be tuned to tell a story after the fact.
flag rate by model
Share of sessions in which the model trips each threshold (seed 0). Darker cells trip more often; the number in each cell is the share, and the median statistic is in the second table.
| Model | high-rate units under-predicted | transients missed | clip-onset failure | rare-pattern failure | rollout divergence | correlation collapse | covariance shrinkage | over-smoothing |
|---|---|---|---|---|---|---|---|---|
| Persistence | 0% | 0% | 50% | 0% | 0% | 0% | 0% | 0% |
| VAR (ridge) | 0% | 13% | 63% | 0% | 0% | 0% | 25% | 100% |
| PCA + VAR | 0% | 13% | 75% | 0% | 0% | 0% | 25% | 100% |
| Smoothed persistence | 0% | 38% | 50% | 0% | 0% | 0% | 25% | 100% |
| LSTM | 0% | 38% | 50% | 0% | 0% | 0% | 38% | 100% |
| Transformer | 0% | 50% | 63% | 0% | 0% | 0% | 25% | 100% |
| LDS (Kalman) | 0% | 38% | 50% | 0% | 0% | 13% | 50% | 100% |
| RNN | 0% | 63% | 38% | 0% | 0% | 0% | 63% | 100% |
| GRU | 0% | 75% | 63% | 0% | 0% | 0% | 38% | 100% |
| TCN | 0% | 50% | 50% | 0% | 88% | 0% | 13% | 100% |
| AR (per unit) | 0% | 100% | 50% | 0% | 0% | 0% | 88% | 100% |
the definitions
- high-rate units under-predicted
- Median relative bias of the forecast mean among the top quarter of units by firing rate. Flagged below −10%.
- transients missed
- Share of the population-count excursion captured in the 2% of test bins with the most total activity. Flagged below 0.5.
- clip-onset failure
- R² in the first 0.5 s after a movie clip restarts, subtracted from R² elsewhere. Flagged above 0.05.
- rare-pattern failure
- R² on bins outside the 95th percentile of the training cloud (10-PC Mahalanobis), subtracted from R² on the rest. Flagged above 0.05.
- rollout divergence
- Lowest free-rollout R² at any horizon up to 2.5 s. Flagged below −1, where the forecast's squared error is more than twice the mean-rate predictor's.
- correlation collapse
- Similarity of the predicted and recorded pairwise-correlation matrices. Flagged below 0.5.
- covariance shrinkage
- Slope of predicted on recorded off-diagonal covariance. Flagged below 0.5.
- over-smoothing
- Predicted over recorded power at 4–10 Hz. Some loss is expected of any rate forecast; flagged only below 0.25.
| Model | high-rate units under-predicted | transients missed | clip-onset failure | rare-pattern failure | rollout divergence | correlation collapse | covariance shrinkage | over-smoothing |
|---|---|---|---|---|---|---|---|---|
| Persistence | 0.000 | +0.688 | +0.047 | −0.333 | −0.841 | +1.000 | +1.000 | +1.000 |
| VAR (ridge) | +0.010 | +0.539 | +0.059 | −0.066 | +0.013 | +0.790 | +0.548 | +0.061 |
| PCA + VAR | +0.013 | +0.570 | +0.089 | −0.053 | +0.024 | +0.718 | +0.557 | +0.026 |
| Smoothed persistence | +0.000 | +0.509 | +0.058 | −0.053 | −0.062 | +0.908 | +0.654 | +0.004 |
| LSTM | −0.006 | +0.504 | +0.049 | −0.060 | −0.098 | +0.698 | +0.566 | +0.013 |
| Transformer | −0.012 | +0.498 | +0.061 | −0.068 | −0.094 | +0.744 | +0.566 | +0.058 |
| LDS (Kalman) | +0.042 | +0.536 | +0.056 | −0.049 | +0.005 | +0.606 | +0.519 | +0.033 |
| RNN | +0.007 | +0.483 | +0.049 | −0.059 | −0.155 | +0.714 | +0.466 | +0.031 |
| GRU | −0.008 | +0.492 | +0.055 | −0.065 | −0.138 | +0.729 | +0.567 | +0.017 |
| TCN | +0.002 | +0.490 | +0.052 | −0.068 | −1.412 | +0.730 | +0.539 | +0.063 |
| AR (per unit) | −0.015 | +0.418 | +0.048 | −0.038 | +0.033 | +0.910 | +0.451 | +0.027 |