AVATAAR通过迭代推理提升长视频问答准确率
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- 分阶段整合全局与局部视频上下文,动态调整检索策略
- 在CinePile上各项任务提升5.6%~8.2%,尤其擅长叙事理解
- 适合需要可解释性长视频分析的场景,如影视解析、教育评测
随着视频内容日益普及,对长视频进行有效理解并回答复杂问题已成为众多应用的关键。尽管大视觉语言模型(LVLMs)已显著提升性能,但在处理需要全面理解和细节分析的细微问题时仍面临挑战。为此,我们提出AVATAAR——一种模块化且可解释的框架,融合全局与局部视频上下文,并引入预检索思维代理与重思模块。AVATAAR构建持久的全局摘要,并在重思模块与预检索思维代理间建立反馈回路,使系统能根据部分答案优化检索策略,实现类人迭代推理。在CinePile基准测试中,相较于基线模型,AVATAAR在时间推理上提升+5.6%,技术类问题+5%,主题类问题+8%,叙事理解+8.2%。实验表明各模块均贡献正向效果,反馈回路是适应性的关键。结果证明,AVATAAR在提升视频理解能力方面表现优异,为长视频问答提供了一种兼具准确性、可解释性与可扩展性的解决方案。
原文摘要 · Abstract (English)
With the increasing prevalence of video content, effectively understanding and answering questions about long form videos has become essential for numerous applications. Although large vision language models (LVLMs) have enhanced performance, they often face challenges with nuanced queries that demand both a comprehensive understanding and detailed analysis. To overcome these obstacles, we introduce AVATAAR, a modular and interpretable framework that combines global and local video context, along with a Pre Retrieval Thinking Agent and a Rethink Module. AVATAAR creates a persistent global summary and establishes a feedback loop between the Rethink Module and the Pre Retrieval Thinking Agent, allowing the system to refine its retrieval strategies based on partial answers and replicate human-like iterative reasoning. On the CinePile benchmark, AVATAAR demonstrates significant improvements over a baseline, achieving relative gains of +5.6% in temporal reasoning, +5% in technical queries, +8% in theme-based questions, and +8.2% in narrative comprehension. Our experiments confirm that each module contributes positively to the overall performance, with the feedback loop being crucial for adaptability. These findings highlight AVATAAR's effectiveness in enhancing video understanding capabilities. Ultimately, AVATAAR presents a scalable solution for long-form Video Question Answering (QA), merging accuracy, interpretability, and extensibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。