Brier Score
Mean squared error of probabilistic forecasts — lower is better calibrated
The Brier score measures how well probabilistic forecasts match outcomes. For binary events, it is the mean squared difference between predicted probability and the realised outcome (0 or 1). Lower is better — a perfect forecast earns 0; confident wrong forecasts score near 1.
Glenn W. Brier introduced the measure in 1950 as a proper scoring rule: it rewards honest reported probabilities, not just lucky guesses. That makes it suitable for track records — comparing forecasters, models, or reviewers on whether their stated confidence matched reality over time.
Role in this corpus
Brier scores operationalise part of authority weight — specifically track record: how past commitments and forecasts in a domain turned out relative to the confidence expressed at bind time.
Strengthening the Commitment Boundary names Brier scores among sources for pre-committed meritocratic weights on collective panels. Weights derived from calibration must be fixed before the vote from prior performance — Brier history, domain accuracy, or validated calibration curves — not adjusted after outcomes are visible. Post-hoc weighting is compliance theatre, not meritocracy.
Article 19 Rule 3 and blind-review override rates are sibling signals: override rate diagnoses decorative human review; Brier score diagnoses whether stated probabilities were honest over time.
How to read it
Brier score does not replace tier discipline or epistemic uncertainty characterisation. It scores probabilistic claims against outcomes, not whether the right capture mode or commitment boundary was used.
Related mechanisms
- Validated calibration curves — complementary; Brier is a scalar summary; curves show where miscalibration lives (e.g. always overconfident at 90%).
- Override-rate tracking — human-in-the-loop analogue for non-probabilistic approve/reject workflows.
- Meritocratic panels — Article 20; Brier-backed weights must appear in the L0 record with their basis.
Corpus stance
A2 — working context: Standard forecast-evaluation metric from meteorology and decision analysis; cited here for commitment-quality and authority-weight design, not as epistemic canon. Use alongside domain-specific accuracy where outcomes are not naturally probabilistic.