Methodology
How we measure our own predictions
Every number on this page comes from a frozen holdout, a named build, and a baseline it had to beat.
Latest scorecard
The first public scorecard publishes after the September 2026 run. Until then this page explains exactly what it will contain.
What each number means
MAPE
Drop timing hit rate
Average view duration error
End hold error
R squared
Spearman rho
95% confidence interval
The baselines the model has to beat
The last three are reported by the training harness. They will appear here with numbers once a trained model is scored on this holdout.
This channel's usual shape
Global median
Position only
Loudness only
The temporal position control
The temporal position control permutes the mapping between modeled and measured curves across videos, keeping each point in its original position. A model that only learned attention decays over time scores the same either way.
A published attention probe reached a correlation of 0.47 and collapsed under this exact control. The number looked real until position alone explained it.
A trained model ships only when its lift over the baseline is positive. Its 95% confidence interval must also exclude zero on this holdout.
Measured versus modeled
Measured
- History bands pulled directly from the platform
- Published audience retention for the exact video
- Fix outcomes: what happened after a creator re-screened a cut
- Viewer panel response, when a panel has run
Modeled
- The modeled retention curve before publish
- Modeled drops and where they land
- Modeled attention moments inside the video
- Script lines the model flagged
- The niche cohort a channel is compared against
