用可审计的LLM融合足球情境,提升比分预测准确性。
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking
- 构建可追踪推理路径的混合架构,结合统计模型与LLM上下文理解
- 最终版本在英超前150场预测中达14.7%准确率,比基线提升4.7%
- 适合关注足球比赛建模、可解释性与概率校准的研究者
足球比分预测兼具强统计基础与复杂情境挑战。动态泊松族模型可估算球队实力、预期进球及合理比分概率,但难以捕捉角色、战术对位、动机变化或首粒进球带来的行为转变。大语言模型(LLM)能推理这些概念,却缺乏概率校准能力。本文通过可审计的信息框架整合两者。提出四次迭代:V1为基于动态得分的Dixon-Coles基准;V2将LLM的上下文评分映射回预期进球参数;V3以冻结的比分候选集进行逐球模拟替代标量修正;V4引入共享首球突破与赛后连锁反应判断、时间感知终止机制及确定性尾部候选。该框架定义输入语义,提供赛前证据,并约束LLM在可检视的推理路径中运行。在2025-26赛季英超前150场比赛的时序回放测试中,V1实现10.0% Top-1与26.7% Top-3精确比分准确率;V3达12.0%与30.0%;V4达14.7%与30.7%。V4将候选覆盖率从77.3%提升至84.7%,但新增尾部候选未进入Top-3。V1原生1X2分布达53.3% argmax准确率,0.9878 log loss,0.5870 Brier score,0.2095 ranked probability score。结果为探索性:训练片段非独立验证集,且闭合式LLM可能存留结果记忆。贡献在于提出可审计的混合架构、清晰的设计演进路径,以及揭示足球感知模拟在何处有效、何处无效的负向发现。
原文摘要 · Abstract (English)
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ratings back into expected-goal parameters; V3, which replaces scalar correction with goal-by-goal simulations over a frozen score-candidate set; and V4, which adds shared first-breakthrough and post-goal cascade judgments, time-aware stopping, and deterministic tail candidates. The harness defines input semantics, supplies pre-match evidence, and constrains the LLM to an inspectable reasoning route. On a chronological replay of the first 150 matches of the 2025-26 English Premier League, V1 achieved 10.0% Top-1 and 26.7% Top-3 exact-score accuracy. V3 reached 12.0% and 30.0%, while V4 reached 14.7% and 30.7%. V4 increased candidate coverage from 77.3% to 84.7%, although no added tail candidate became a Top-3 exact hit. V1's native 1X2 distribution achieved 53.3% argmax accuracy, 0.9878 log loss, 0.5870 Brier score, and 0.2095 ranked probability score. These results are exploratory: the development slice is not an untouched benchmark, and temporal input isolation cannot exclude outcome memory in a closed LLM. The contribution is an auditable hybrid architecture, a clear design evolution, and negative findings showing where football-aware simulation does and does not improve score selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。