HB Betting Experiment

Model Health

Feed uptime, guardrail effectiveness and probability calibration. These metrics decide when the experiment may advance past paper validation.

Active validation stagePAPER VALIDATIONReal-money staking stays disabled until calibration and CLV thresholds hold across a full season.

Odds API uptime

96.9%

OK

Kalshi uptime

85.1%

OK

Storage uptime

100.0%

OK

Degraded / failed runs

76

Last run —

Stale-data alerts

172

Runs with odds older than 900s

Zero-liquidity rejections

176

Guardrail blocked untradeable markets

Market-match preventions

15

Incorrect market pairings blocked pre-bet

Brier score

0.230

Lower is better · 0.25 = coin flip

Rejection reasons

Why candidates were passed or declined.

  • Edge below threshold478
  • Bookmaker count below minimum (< 4)335
  • Zero liquidity on Kalshi market176
  • Commence time inside blackout window17
  • Stale odds (> 900s)16
  • Market match confidence too low15
  • Consensus dispersion too high13

Bookmaker count distribution

How many books priced each candidate. Fewer than 4 is auto-rejected.

Consensus dispersion

Standard deviation across book prices. High dispersion means an unreliable consensus.

Model calibration

Predicted probability vs. observed win rate per bucket.

BucketPredictedObservedNGap
0–10%5%0%2-5%
10–20%15%27%11+12%
20–30%25%23%13-2%
30–40%35%49%49+14%
40–50%45%54%71+9%
50–60%55%64%83+9%
60–70%65%72%95+7%
70–80%75%56%45-19%
80–90%85%87%15+2%
90–100%95%80%10-15%

Stale-data alerts

Runs where the odds snapshot exceeded the 900-second freshness budget.

  • run_00111127s oldSYSTEM OK
  • run_00321238s oldSYSTEM OK
  • run_00401339s oldSYSTEM DEGRADED
  • run_0052961s oldSYSTEM OK
  • run_0072956s oldSYSTEM OK
  • run_00821389s oldSYSTEM OK
  • run_00901369s oldSYSTEM OK
  • run_01111273s oldSYSTEM OK
  • run_01211249s oldSYSTEM OK
  • run_01401253s oldSYSTEM OK
  • run_0151902s oldSYSTEM OK
  • run_0161901s oldSYSTEM OK