arXiv:2511.16901cs.CV2025-11AAAI被引 1

构建首个真实场景音视频时空推理数据集,推动视频大模型理解复杂事件。

R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios

  • 设计多阶段流程生成细粒度音视频时空标注数据
  • 覆盖5000+未剪辑视频,含2.7万物体与100类事件
  • 提出无中间监督的强化学习模型,提升复杂场景推理能力

近年来,多模态大语言模型在视频理解任务上取得快速进展,但现有研究多聚焦于简单视频场景,难以反映真实世界音视频事件的复杂多样性。为此,我们首次提出R-AVST数据集,专为音视频细粒度时空推理设计,包含基于大模型的关键对象提取、自动空间标注与人工质量检查的构建流水线,涵盖超过5000段未剪辑视频,涉及27000个物体和100类音视频事件。在此基础上,定义三类核心时空推理任务,并生成超过8000个高质量、分布均衡的问答对,有效评估模型性能。为增强推理能力,我们提出AVST-Zero,一种基于强化学习的模型,无需中间监督,通过精心设计的多维奖励直接优化行为。大量实验验证了R-AVST在推进音视频时空推理方面的有效性,且AVST-Zero表现优于现有模型。据我们所知,R-AVST是首个面向真实世界音视频时空推理的专用数据集,而AVST-Zero为该领域未来挑战提供了新思路。

原文摘要 · Abstract (English)

Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video scenarios, failing to reflect the complex and diverse nature of real-world audio-visual events in videos. To bridge this gap, we firstly introduce R-AVST, a dataset for audio-visual reasoning featuring fine-grained spatio-temporal annotations. In constructing this, we design a pipeline consisting of LLM-based key object extraction, automatic spatial annotation and manual quality inspection, resulting in over 5K untrimmed videos with 27K objects across 100 types of audio-visual events. Building on this dataset, we define three core tasks for spatio-temporal reasoning in audio-visual scenes and generate more than 8K high-quality, evenly distributed question-answer pairs to effectively benchmark model performance. To further enhance reasoning, we propose AVST-Zero, a reinforcement learning-based model that avoids intermediate supervision, directly optimizing behavior via carefully designed multi-dimensional rewards. Extensive experiments validate the effectiveness of our R-AVST in advancing audio-visual spatio-temporal reasoning, upon which AVST-Zero demonstrates competitive performance compared to existing models. To the best of our knowledge, R-AVST is the first dataset designed for real-world audio-visual spatio-temporal reasoning, and AVST-Zero offers a novel perspective for tackling future challenges in this domain.

音视频理解时空推理多模态大模型数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。