用足球场景测试视觉语言模型的战术决策能力,发现其决策保守且价值判断失准。
SportD: How do VLMs physically strategize?

- 构建足球决策数据集SportD,评估模型在真实比赛场景中的策略选择。
- 模型选对最优动作仅30%正确率,远低于人类,且偏好低风险低回报行动。
- 揭示模型将成功概率误当作价值,导致风险规避偏差,适合研究模型决策机制者关注。
视觉语言模型(VLMs)能描述场景,但能否在其中做出良好决策?我们以足球为测试基准,通过可量化的动作评估其战略决策能力。引入SportD数据集,包含1415个来自职业男女足球比赛的决策场景,模型需观察决策前数秒画面并选择下一步动作。结果显示,模型仅约30%时间选择最优动作,甚至低于人类表现。此外,模型明显倾向安全动作,偏好低方差、低收益的选择,且进展更少。前沿模型虽能较好估计动作成功率(83%-92%将高成功率动作列于前列),但系统性混淆成功概率与价值:赋予更可能成功的动作更高价值(ρ=+0.30至+0.52),而真实数据中二者相关性极弱(ρ=-0.08)。这种保守源于价值判断失准。通过替换一句推理语句,引导模型承担风险,可显著提升性能。SportD为评估VLM物理策略决策提供了新范式,揭示了分解决策过程可揭示系统性偏差如风险规避的内在机制。
原文摘要 · Abstract (English)
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action. Models only select the optimal action around 30% of the time, even less frequently than humans do. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 83-92% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed ($ρ$=+0.30 to +0.52), despite no such relationship in the ground truth ($ρ$=-0.08). The conservatism therefore reflects a mis-calibration of value. Replacing a single deliberation sentence with one that steers toward risk lifts the frontier models towards the real players' skills. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。