arXiv:2511.00370cs.CVcs.AI2025-11

用多智能体冲突机制提升视频片段检索的准确性和可信度

Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict

  • 引入强化学习与多智能体框架,统一定位边界与证据生成
  • 通过冲突分析实现无需额外训练的越界查询识别
  • 适合需要高可靠性、抗干扰能力的现实视频交互场景

视频片段检索任务旨在通过文本查询从长视频中定位对应时间段。现有方法未考虑不同模型输出间的冲突,导致结果难以有效融合。本文提出一种基于强化学习的单次扫描检索模型,可同时确定片段边界并生成定位证据。进一步设计多智能体系统,利用证据学习解决各智能体间定位结果的冲突。作为副产物,该机制可无需额外训练即判断查询是否超出视频范围,契合真实应用需求。在多个基准数据集上的实验表明,所提方法优于当前最优方案。研究还发现,建模多智能体间的竞争与冲突是提升强化学习性能的有效路径,并揭示了证据学习在多智能体框架中的新作用。

原文摘要 · Abstract (English)

Video moment retrieval uses a text query to locate a moment from a given untrimmed video reference. Locating corresponding video moments with text queries helps people interact with videos efficiently. Current solutions for this task have not considered conflict within location results from different models, so various models cannot integrate correctly to produce better results. This study introduces a reinforcement learning-based video moment retrieval model that can scan the whole video once to find the moment's boundary while producing its locational evidence. Moreover, we proposed a multi-agent system framework that can use evidential learning to resolve conflicts between agents' localization output. As a side product of observing and dealing with conflicts between agents, we can decide whether a query has no corresponding moment in a video (out-of-scope) without additional training, which is suitable for real-world applications. Extensive experiments on benchmark datasets show the effectiveness of our proposed methods compared with state-of-the-art approaches. Furthermore, the results of our study reveal that modeling competition and conflict of the multi-agent system is an effective way to improve RL performance in moment retrieval and show the new role of evidential learning in the multi-agent framework.

视频检索多智能体强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。