Pitwall实时生成F1策略简报,确保每句话都有数据支撑。
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
- 用蒙特卡洛模拟+验证器确保每句简报符合实时比赛状态
- 90.3%预测准确率,155次回测中胜者进前三,误差仅0.0745
- 支持英/西/葡三语,适合体育直播与赛事分析场景
实时体育解说是在时间压力下的真实内容生成:语句涉及具体运动员,状态每几秒变化,且生成时无参考文本。我们提出Pitwall,一个生产级系统,能以英语、西班牙语和葡萄牙语生成实时的F1策略简报,将忠实性作为架构属性而非目标:每个发布句子被分解为类型化事实主张(排名、差距、轮胎、速度、超车、赛事控制),并由其触发的概率化比赛状态进行验证。同一验证器也用于筛选微调数据:在3,045条模型生成的目标中,仅81.9%的每项主张均获状态支持,其余退回可证明忠实的模板,使生成器从未接触无依据目标。验证有效源于底层基础:基于126场比赛(2018–2024)校准、在完全未见的2025–2026赛季验证的向量化蒙特卡洛引擎(每圈2,000次延续),在155次回测中,获胜者进入前三的准确率达90.3%,未见样本的布里尔分数为0.0745。系统两部分均发现:优势存在权衡,必须分别管控。仿真显示,校准最优并非决策最优;生成中,对更丰富目标的微调带来生动性,但在状态稀疏时引发幻觉——四基线复现证实此问题源于基础模型指令遵循,非规模问题,而稀疏上下文审计可消除该缺陷。端到端流程(实时计时至已验证的多语言简报)已在2026年奥地利与英国大奖赛连续验证;在银石赛道,一时间戳概率轨迹在结果揭晓前已写入磁盘,十圈前便锁定最终冠军。
原文摘要 · Abstract (English)
Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the grounding state changes every few seconds, and no reference text exists at generation time. We present Pitwall, a production system that generates natural-language Formula 1 strategy briefings in English, Spanish, and Portuguese, treating faithfulness as an architectural property rather than an aspiration: every published sentence is decomposed into typed factual claims (positions, gaps, tyres, pace, overtakes, race control) and each claim is verified against the probabilistic race state that prompted it. The same verifier gates the fine-tuning data: of 3,045 model-written targets, only the 81.9% whose every claim is state-supported are retained, the rest falling back to a provably faithful template, so the generator never sees an ungrounded target. Verification is meaningful because of the grounding substrate: a vectorized Monte Carlo engine (N=2,000 per-lap race continuations) calibrated on 126 races (2018-2024) and validated on fully held-out 2025-2026 seasons (winner-in-top-3 90.3% over 155 backtests; held-out Brier 0.0745). A recurring finding spans both halves of the system: virtues trade off and must be gated separately. In simulation, calibration-optimal is not decision-optimal; in generation, fine-tuning on richer targets buys vividness that collapses into hallucination when the grounding state is sparse -- a failure a four-base replication traces to base-model instruction adherence, not scale, and that sparse-context auditing removes from the production model. End-to-end operation -- live timing to verified trilingual briefings -- was confirmed at two consecutive live Grands Prix (Austria and Britain, 2026); at Silverstone a timestamped probability trace, committed to disk before the outcome was known, locked onto the eventual winner ten laps before the flag.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。