arXiv:2606.03635cs.CVcs.AI2026-06

构建视频隐含信息理解基准,评估模型对短视频深层意图的推理能力

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

论文配图:VidMsg: A Benchmark for Implicit Message Inference in Short Videos
图 1 · 摘自论文原文
  • 基于消息优先流程生成400个隐含意图视频片段
  • 现有模型在隐含信息识别上表现不佳,需整合上下文线索
  • 适合视频搜索、推荐系统等需要深层理解的应用场景

理解短时在线视频不仅需识别可见对象和动作,还需捕捉创作者隐藏的深层意图。我们提出VidMsg,一个面向互联网原生短视频的隐含信息理解基准。该数据集包含400个来自YouTube的视频片段,覆盖9个实际主题领域和52个细粒度目标信息,涵盖职业与金融、教育、健康与福祉、文化、安全、可持续发展及生活方式等。数据集通过‘消息优先’流程构建:先由大语言模型将目标信息转化为间接搜索场景,再检索候选视频,最后由人工标注保留能传达预期信息但不直白的片段。VidMsg旨在支持双向消息-视频检索,适用于视频搜索与推荐等规模化应用,要求系统具备整体理解能力。除检索任务外,还提供诊断性多选问答基准,测试模型从语义相近选项中选出真实意图的能力。实验表明,当前主流视频-语言与检索模型在该基准上表现不佳,因任务需语用推理、上下文整合及语义相近信息的区分。我们还提出了基线方法VidVec-Msg,提升消息导向检索效果,但仍留有巨大改进空间。

原文摘要 · Abstract (English)

Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purpose in the clip. We introduce VidMsg, a benchmark for evaluating implicit message understanding in short, internet-native video clips. VidMsg contains 400 YouTube-derived clips across 9 practical topic areas and 52 fine-grained target messages, covering domains such as career and finance, education, health and well-being, culture, safety, sustainability, and lifestyle. VidMsg is constructed through a message-first pipeline: an LLM first translates target messages into indirect search scenarios, which are used to retrieve candidate clips. Human annotators then retain clips that convey the intended message without being overly explicit. VidMsg is designed primarily for bidirectional message-clip retrieval for scalable applications such as video search and recommendation, where systems must capture holistic video understanding. In addition to retrieval, VidMsg includes a diagnostic multiple-choice QA benchmark, where models select the intended message of a clip from semantically related alternatives. Experiments with contemporary video-language and retrieval models show that strong models often fail on VidMsg, because the task requires pragmatic inference, integration of contextual cues, and discrimination among semantically close messages. We also introduce VidVec-Msg, a baseline method that improves message-oriented retrieval while leaving substantial headroom for future work.

视频理解隐含意图基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。