Coherence training for forecasts · Kelvin, MATS

Now: 32B baseline → 0.6B

Qwen3-32B first: same family and training cutoff as the 0.6B student.

Target

0.50.550.60.650.7Constant 0.50.693Form-only prior0.607Base-rate table0.578Qwen3-32B zero-shot0.616Qwen3-0.6B + consistencynexttest log loss, lower is better
Test set: 1,066 pairs of real Manifold questions (May–Dec 2025), 12 logical forms each. A model must beat the 12-number base-rate table.

Next

Baseline result

Done

Qwen3-32B, zero-shot, 23 September 2026. 1,066 test pairs, 12 forms each, scored once. Prompt chosen on validation (6 variants). No unreadable answers.

PredictorLog loss, observed formLog loss, all 12 formsAccuracyECE
Constant 0.50.6930.693
Form-only prior0.6070.606
Base-rate table0.5780.570
Qwen3-32B0.6160.6110.7030.062
Qwen3-32B, calibrated on validation0.5970.5950.7030.037

32B does not beat the base-rate table. Weakest on negations: P and not-P differ from summing to 1 by more than 0.2 in 35% of pairs.

000.250.250.50.50.750.7511803 answers, said 0.05, happened 0.152034 answers, said 0.15, happened 0.221902 answers, said 0.25, happened 0.341169 answers, said 0.35, happened 0.41450 answers, said 0.44, happened 0.42302 answers, said 0.54, happened 0.532071 answers, said 0.65, happened 0.642388 answers, said 0.75, happened 0.70719 answers, said 0.85, happened 0.76954 answers, said 0.95, happened 0.85stated probabilityhow often it happened
Below the diagonal at low probabilities: the 32B model is too confident that things will not happen. ECE 0.062.

Frontier traces remember, not forecast

00.050.10.15Polymarket price0.128Frontier model (Sol)0.111Brier, pre-2025 questions
On pre-2025 questions a frontier model beats the market price: it recalls outcomes. So its traces are not used for training, and all evaluation is post-cutoff.