arXiv:2602.03707cs.CL2026-02被引 1

低成本长视频问答新方法,能智能检索音视频片段并自主规划回答。

OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering

  • 构建代理循环,动态调用工具、规划步骤并融合多模态证据。
  • 在低资源下实现比现有方法更优的准确率,长视频问答性能显著提升。
  • 适合需要高效处理长时序音视频数据的研究与应用,如智能客服、教育系统。

长时程多模态问答需基于文本、图像、音频和视频进行推理。尽管近期在通用多模态大模型(OmniLLMs)方面取得进展,低资源长音视频问答仍面临密集编码成本高、细粒度检索弱、主动规划能力有限及缺乏端到端优化等问题。为此,我们提出 OmniRAG-Agent,一种面向预算受限的长音视频推理的智能体式多模态问答方法。该方法构建图像-音频检索增强生成模块,使 OmniLLM 能从外部知识库中获取短且相关的帧与音频片段;同时引入代理循环机制,实现跨轮次的规划、工具调用与检索证据融合;进一步采用组相对策略优化(group relative policy optimization),联合提升工具使用效率与答案质量。在 OmniVideoBench、WorldSense 与 Daily-Omni 数据集上的实验表明,OmniRAG-Agent 在低资源设置下持续优于先前方法,并取得强性能表现,消融实验证明各组件有效性。

原文摘要 · Abstract (English)

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding, weak fine-grained retrieval, limited proactive planning, and no clear end-to-end optimization. To address these issues, we propose OmniRAG-Agent, an agentic omnimodal QA method for budgeted long audio-video reasoning. It builds an image-audio retrieval-augmented generation module that lets an OmniLLM fetch short, relevant frames and audio snippets from external banks. Moreover, it uses an agent loop that plans, calls tools across turns, and merges retrieved evidence to answer complex queries. Furthermore, we apply group relative policy optimization to jointly improve tool use and answer quality over time. Experiments on OmniVideoBench, WorldSense, and Daily-Omni show that OmniRAG-Agent consistently outperforms prior methods under low-resource settings and achieves strong results, with ablations validating each component.

多模态长视频智能体问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。