arXiv:2602.19040cs.IRcs.AI2026-02

动态协作多智能体系统提升复杂文本到视频检索效果

Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval

  • 按查询需求动态调度检索、推理、重写三类智能体协同工作
  • 在三个TRECVid数据集上性能比CLIP4Clip提升一倍
  • 适合需要精准时序推理的视频检索场景

短视频平台兴起与多模态大模型的发展,推动了可扩展、高效的零样本文本到视频检索系统的需求。尽管大规模预训练提升了跨模态对齐能力,现有方法仍难以处理依赖查询的时序推理,限制了其在包含时序、逻辑或因果关系的复杂查询上的表现。为此,我们提出一种自适应多智能体检索框架,根据每个查询的需求,在多轮推理中动态协调专业化智能体。框架包含:(1) 可扩展检索大型视频语料库的检索智能体;(2) 零样本上下文时序推理的推理智能体;(3) 重写模糊查询并恢复迭代中性能下降的查询重写智能体。这些智能体由编排智能体动态协调,利用中间反馈和推理结果指导执行。我们还引入一种新通信机制,融合检索性能记忆和历史推理轨迹以增强协调与决策。在覆盖八年数据的三个TRECVid基准上实验表明,该框架性能是CLIP4Clip的两倍,并显著超越当前最优方法。

原文摘要 · Abstract (English)

The rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrieval systems. While recent advances in large-scale pretraining have improved zero-shot cross-modal alignment, existing methods still struggle with query-dependent temporal reasoning, limiting their effectiveness on complex queries involving temporal, logical, or causal relationships. To address these limitations, we propose an adaptive multi-agent retrieval framework that dynamically orchestrates specialized agents over multiple reasoning iterations based on the demands of each query. The framework includes: (1) a retrieval agent for scalable retrieval over large video corpora, (2) a reasoning agent for zero-shot contextual temporal reasoning, and (3) a query reformulation agent for refining ambiguous queries and recovering performance for those that degrade over iterations. These agents are dynamically coordinated by an orchestration agent, which leverages intermediate feedback and reasoning outcomes to guide execution. We also introduce a novel communication mechanism that incorporates retrieval-performance memory and historical reasoning traces to improve coordination and decision-making. Experiments on three TRECVid benchmarks spanning eight years show that our framework achieves a twofold improvement over CLIP4Clip and significantly outperforms state-of-the-art methods by a large margin.

文本到视频多智能体时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。