用熵驱动策略自动选择提问或优化,提升视频搜索精准度
ADEPT: An Entropy-Driven Dual-Strategy Agent for Interactive Video Retrieval

- 基于熵值动态选择提问或优化查询,无需训练
- 在两个数据集上超越现有非交互式与视频大模型方法
- 适合需要精准视频检索的交互式系统开发者
针对海量视频数据中用户查询模糊导致的检索难题,现有单轮检索范式因缺乏有效反馈机制而性能受限。根源在于‘意图-查询鸿沟’:简单文本无法捕捉用户真实意图。为此,本文提出ADEPT框架——一种无需训练的双策略智能体,通过熵驱动决策引擎,在ASK(提问)与REFINE(优化)策略间动态切换。实验在两个挑战性数据集上验证,ADEPT显著优于所有非交互式、启发式及视频大模型基线。核心贡献在于构建了一种高效可解释的交互式检索策略,为该领域设立新基准。
原文摘要 · Abstract (English)
This research aims to solve the challenge of video retrieval from massive datasets, caused by ambiguous user queries. Prevailing single-round retrieval paradigms face a performance bottleneck, as they lack effective feedback mechanisms to handle complex search intentions. The root cause is the "Intent-Query Gap", where users' intent cannot be captured by a simple text query. To solve this, we propose the ADEPT framework: a training-free agent that pioneers an entropy-driven decision engine to efficiently guide dialogue by dynamically selecting between ASK and REFINE strategies. Experiments on two challenging datasets demonstrate that ADEPT significantly outperforms all non-interactive, heuristic, and Video-LLM baselines. The core contribution of this work is an efficient and interpretable entropy-driven interactive strategy that sets a new performance benchmark for the field of interactive video retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。