用三阶段流程提升视频检索生成的准确性和连贯性。
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation

- 分三阶段处理:查询扩展融合、结构化证据提取、可控生成
- 检索性能nDCG@10从0.195提升至0.759,生成人类评分达3.83
- 适合需要多视频推理与高引用召回的生成任务
从视频中进行检索增强生成,需从大规模语料库中检索相关音视频证据,并将其合成连贯且可追溯的文本。现有方法在两端均存在瓶颈:检索方法难以应对复杂多面的查询,无法通过单一嵌入捕捉;生成方法缺乏跨多视频的高层推理能力,且在长视频上下文下受限于记忆容量。我们提出MARQUIS:一个三阶段流水线,分别通过(1)查询扩展、融合与重排序,(2)校准的结构化证据提取,(3)基于提取证据的篇章生成(可由RLM控制)。在MAGMaR2026共享任务中,检索性能从0.195提升至0.759(nDCG@10)。生成方面,ITER-QA-BASE平均人类评分从3.09升至3.83,而MARQUIS-RLM达到人类评分3.30,并在非QA系统中实现最强引用召回率。
原文摘要 · Abstract (English)
Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into coherent, attributed text. Current approaches struggle at both ends: retrieval methods fail on complex, multi-faceted queries that cannot be captured by a single embedding, while generation methods lack the high-level reasoning needed to synthesize across multiple videos and face memory constraints over long, multi-video contexts. We present MARQUIS: a three-stage pipeline that addresses these limitations through (1) query expansion, fusion, and reranking, (2) calibrated structured evidence extraction, and (3) article generation from extracted evidence, optionally controlled by an RLM. On the MAGMaR2026 shared task, we improve retrieval performance from 0.195 to 0.759 (nDCG@10). For article generation, ITER-QA-BASE improves average human score from 3.09 to 3.83 over the CAG baseline, while MARQUIS-RLM achieves a human score of 3.30 and the strongest citation recall among non-QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。