用大模型补全模糊视频查询的上下文,提升检索准确率
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
- 通过大模型自动补充缺失的位置、背景等上下文信息
- 在300小时视频中实现高速高精度事件检索,显著优于基线方法
- 适合处理不完整或模糊的文本查询,如缺少地点描述的搜索
基于文本查询从视频中检索事件正变得愈发困难,因多媒体内容快速增长。现有方法多聚焦物体层面描述,忽视上下文信息的重要性,尤其在查询缺乏位置细节或背景模糊时表现不佳。为此,我们提出RAPID(检索增强并行推理草稿系统),利用大语言模型和提示学习,语义上修正并丰富用户查询的上下文信息。这些增强后的查询通过并行检索处理,并经评估步骤选择与原始查询最匹配的结果。在自建数据集上的大量实验表明,RAPID在上下文不完整的查询上显著优于传统方法。系统在2024年胡志明市人工智能挑战赛中验证,成功从超过300小时视频中检索事件。与比赛组织方提出的基线相比,本方法展现出更强的有效性与鲁棒性。
原文摘要 · Abstract (English)
Retrieving events from videos using text queries has become increasingly challenging due to the rapid growth of multimedia content. Existing methods for text-based video event retrieval often focus heavily on object-level descriptions, overlooking the crucial role of contextual information. This limitation is especially apparent when queries lack sufficient context, such as missing location details or ambiguous background elements. To address these challenges, we propose a novel system called RAPID (Retrieval-Augmented Parallel Inference Drafting), which leverages advancements in Large Language Models (LLMs) and prompt-based learning to semantically correct and enrich user queries with relevant contextual information. These enriched queries are then processed through parallel retrieval, followed by an evaluation step to select the most relevant results based on their alignment with the original query. Through extensive experiments on our custom-developed dataset, we demonstrate that RAPID significantly outperforms traditional retrieval methods, particularly for contextually incomplete queries. Our system was validated for both speed and accuracy through participation in the Ho Chi Minh City AI Challenge 2024, where it successfully retrieved events from over 300 hours of video. Further evaluation comparing RAPID with the baseline proposed by the competition organizers demonstrated its superior effectiveness, highlighting the strength and robustness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。