arXiv:2412.01694cs.CV2024-12CVPR被引 41

通过自动生成思维链提升视频问答模型的推理能力

Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation

  • 用代理系统分解复杂问题,生成视觉模型处理的中间推理链
  • 在多选和开放题基准上显著提升模型性能
  • 适合需要可解释性与时空理解的视频分析研究者

本文针对视频问答(VideoQA)任务中所需的多步推理与时空动态理解问题。尽管大型视频语言模型在基准测试中表现良好,但常缺乏可解释性与时空定位能力。为此,我们提出代理思维链蒸馏(Agent-of-Thoughts Distillation, AoTD),将自动生成的思维链(CoTs)融入指令微调过程。具体而言,利用基于代理的系统将复杂问题分解为子任务,并由专用视觉模型分别处理,中间结果作为推理链;同时引入大语言模型进行验证,确保生成思维链的可靠性。大量实验表明,AoTD在多项选择与开放回答基准上均取得性能提升。

原文摘要 · Abstract (English)

This paper tackles the problem of video question answering (VideoQA), a task that often requires multi-step reasoning and a profound understanding of spatial-temporal dynamics. While large video-language models perform well on benchmarks, they often lack explainability and spatial-temporal grounding. In this paper, we propose Agent-of-Thoughts Distillation (AoTD), a method that enhances models by incorporating automatically generated Chain-of-Thoughts (CoTs) into the instruction-tuning process. Specifically, we leverage an agent-based system to decompose complex questions into sub-tasks, and address them with specialized vision models, the intermediate results are then treated as reasoning chains. We also introduce a verification mechanism using a large language model (LLM) to ensure the reliability of generated CoTs. Extensive experiments demonstrate that AoTD improves the performance on multiple-choice and open-ended benchmarks.

视频问答思维链推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。