Evaluation
Ranked probability score on ordinal buckets, log loss, and interval coverage. 2-day and 7-day are scored separately. If the fancy models lose to linear and M1, we say so and ship the baseline. This is opt-in — the harness is slow on purpose.
Train on same-horizon windows 1…k, score window k+1 at f = 0.1, 0.25, 0.5, 0.75, 0.9. Features at t use only posts with created_at ≤ t. Press the button when you can wait.