Prospective prediction systems should not only generate individual-level predictions but also indicate when event-level conditions suggest high upset risk or difficult-to-act-on cases. This study developed and evaluated a leakage-aware race-level upset-risk diagnostic for prospective horse-race prediction under temporal validation. Japanese flat-racing data were analyzed using a fixed pre-event evaluation framework. A race-level binary upset label was constructed retrospectively from post-event results, payouts, and final popularity for label construction and evaluation only, whereas prediction features were restricted to pre-event race-structure variables. The temporal split used 2015–2022 races for training, 2023–2024 races for validation, and races from January 5, 2025 to May 10, 2026 for independent testing. Candidate-model assessment was performed using the validation split only, and the primary revised model was a prespecified unweighted standardized logistic regression model. In the temporal test set of 4,556 races, the primary revised model achieved ROC AUC: 0.6459, PR-AUC: 0.6308, Brier score: 0.2325, and log loss: 0.6567. Upset-risk stratification showed that the top 10% highest-upset-probability races had an observed upset rate of 0.6820, compared with a baseline rate of 0.5151. Conversely, the bottom 10% lowest-upset-probability races had an observed upset rate of 0.2193. These findings suggest that a race-level upset-risk diagnostic can stratify races into high-risk and lower-risk strata while remaining separate from horse-level prediction scores. The proposed framework emphasizes prediction-time information constraints, leakage prevention, validation-only candidate-model assessment, calibration assessment, and selective-use evaluation as practical components of deployment-oriented machine-learning evaluation.
Leakage-aware race-level upset-risk diagnostics for prospective horse-race prediction
Shuichi Sugiura

