arXiv:2608.09435cs.AI2026-08

让模型同时听声、看景、追踪声音移动,实现多模态动态声音理解。

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

论文配图:Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
图 1 · 摘自论文原文
  • 融合全景视觉与空间音频,构建声音定位与轨迹推理机制。
  • 在40万题数据集上达到77.83%平均准确率,超越基线一倍以上。
  • 适用于需要时空声音感知的智能监控、机器人导航场景。

理解动态声源需同时确定声音来源、位置及随时间的运动轨迹。现有音视频模型常将音频片段视为全局事件,而视觉语言模型缺乏空间音频线索以定位和追踪独立声源。为此,我们引入ST-OmniQA——一个基于全景视频与同步一阶全向声(FOA)音频的时空音视频问答基准,包含40,000段视频与400,000个问答对,涵盖四个能力层级:声音事件识别、到达方向、声源距离、运动轨迹及时间锚定的音视频推理。基于此,我们提出ST-Omni-R1,通过渐进式课程学习与推理树强化学习,整合FOA提取的语义与轨迹表征与全景视觉上下文。该模型在四个层级上平均语义准确率达77.83%,远超最优基线37.28%。在三个公开空间音频基准上的结果表明,其学习的空间与运动表征具有良好的迁移能力。

原文摘要 · Abstract (English)

Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

音视频理解空间音频多模态轨迹追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。