arXiv:2604.15736cs.CVcs.CL2026-04

首个评估视频多模态大模型当裁判的基准,发现当前模型准确率仅六成左右。

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

论文配图:RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
图 1 · 摘自论文原文
  • 构建覆盖11项运动的6475个问答数据集,评估判罚、规则推理等五项核心能力。
  • 最强模型准确率约60%,开源模型仅47%,普遍误判正常动作为犯规。
  • 揭示模型在规则应用和时间定位上的短板,适合研究多模态决策与体育AI的读者。

尽管多模态大语言模型在通用视频理解方面表现优异,但其在需要规则约束的专业决策任务中仍缺乏充分探索。本文提出RefereeBench,首个用于评估多模态大模型作为自动体育裁判的大规模基准。该基准涵盖11种运动,包含925段精选视频和6,475个问答对,评估五项核心判罚能力:犯规存在性判断、犯规与处罚分类、规则推理、实体感知和时间定位。所有标注均经人工完成,确保符合真实判罚逻辑与多模态证据。对现有先进模型的测试显示,即使是最强模型(如Doubao-Seed-1.8和Gemini-3-Pro)准确率也仅约60%,最强开源模型Qwen3-VL仅为47%。结果表明当前模型远未达到可靠裁判水平。进一步分析发现,模型虽能识别事件与涉事主体,但在规则应用和时间定位上表现不佳,常将正常动作误判为犯规。该基准凸显未来模型需更好融合领域知识与多模态理解,推动可信的AI辅助判罚及更广泛的多模态决策发展。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first large-scale benchmark for evaluating MLLMs as automatic sports referees. Spanning 11 sports with 925 curated videos and 6,475 QA pairs, RefereeBench evaluates five core officiating abilities: foul existence, foul and penalty classification, foul and penalty reasoning, entity perception, and temporal grounding. The benchmark is fully human-annotated to ensure high-quality annotations grounded in authentic officiating logic and multimodal evidence. Extensive evaluations of state-of-the-art MLLMs show that even the strongest models, such as Doubao-Seed-1.8 and Gemini-3-Pro, achieve only around 60% accuracy, while the strongest open-source model, Qwen3-VL, reaches only 47%. These results indicate that current models remain far from being reliable sports referees. Further analysis shows that while models can often identify incidents and involved entities, they struggle with rule application and temporal grounding, and frequently over-call fouls on normal clips. Our benchmark highlights the need for future MLLMs that better integrate domain knowledge and multimodal understanding, advancing trustworthy AI-assisted officiating and broader multimodal decision-making.

多模态大模型体育裁判评测基准规则推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。