首个面向中文游戏短视频帧搜索的多模态评测基准,测试模型在复杂场景下的知识检索与推理能力。
SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

- 构建中文游戏领域短视频帧搜索评测集,含5000个四选一测试题
- 实际代理搜索最高准确率79.1%,远低于理想知识水平95.4%
- 揭示视觉定位、检索质量、工具使用等关键瓶颈问题
多模态大语言模型正被用作智能体核心,以理解多模态输入、规划检索动作、调用外部工具并基于检索信息推理。然而现有评测基准极少评估其在短视频应用中的表现,尤其在画面暂停时视觉模糊、需依赖垂直领域、长尾且快速演化的知识背景下。我们提出SVFSearch,首个面向中文游戏领域的短视频帧搜索开放评测基准。该数据集包含5,000个四选一测试样例和4,198个辅助训练样例,每个样本源自真实短视频片段中的定格游戏画面。为确保公平可复现评估,SVFSearch提供冻结的离线检索环境,包含游戏领域文本语料库、话题关联图像图库及文本、图像、多模态检索接口,避免依赖不可控的网络搜索API。我们评估了从直接问答到RAG流程、计划-执行-重规划智能体以及学习型搜索模型等多种范式。结果表明,模型仅回答与实际代理搜索和理想知识之间存在显著差距:最佳开源直接问答模型准确率为66.4%,最佳实用代理达到79.1%,而理想知识水平可达95.4%。进一步分析揭示了视觉定位、检索质量、证据驱动推理与工具使用行为中的多重瓶颈,包括过度搜索、仅答不搜、检索诱发误导等问题。
原文摘要 · Abstract (English)
Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external tools, and reason over retrieved information. Yet existing benchmarks rarely evaluate this ability in short-video applications, where a paused frame is often visually ambiguous and answering requires vertical, long-tail, and fast-evolving domain knowledge. We introduce SVFSearch, the first open benchmark for short-video frame search in the Chinese gaming domain. SVFSearch contains 5,000 four-choice test examples and 4,198 auxiliary training examples, each centered on a paused game scene from a real short-video clip. To support fair and reproducible evaluation, SVFSearch provides a frozen offline retrieval environment with a game-domain text corpus, a topic-linked image gallery, and text, image, and multimodal retrieval interfaces, avoiding reliance on uncontrolled web search APIs. We evaluate representative paradigms ranging from direct QA and RAG workflow to Plan-Act-Replan agents and learned search models. Results reveal a large gap between model-only answering, practical agentic search, and oracle knowledge: the best open-source direct-QA model reaches 66.4%, the best practical agent achieves 79.1%, and oracle knowledge reaches 95.4%. Further analysis exposes bottlenecks in visual grounding, retrieval quality, evidence-grounded reasoning, and tool-use behavior, including over-search, answer-only shortcuts, and retrieval-induced misleading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。