arXiv:2510.24693cs.SDcs.CL2025-10被引 9

构建音频4D智能评测基准,测试声音在时空中的精细推理能力。

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

  • 通过物理仿真与人工标注构建高质量音频数据集
  • 模型在时间与空间推理上准确率下降超30%,凸显语言描述盲区
  • 揭示闭源模型感知瓶颈,开放模型全面落后

尽管多模态大语言模型和大型音频-语言模型进展迅速,现有音频评测大多依赖文本描述可恢复的语义,掩盖了细粒度感知推理的缺陷。本文提出音频4D智能概念,即对时间与三维空间中声音动态的推理能力,并引入STAR-Bench评测基准。该基准包含基础声学感知(六种属性,绝对与相对两种情形)与整体时空推理(连续与离散过程的片段重排,以及静态定位、多源关系、动态轨迹等空间任务)。数据构建采用双重方法:基础任务使用程序生成与物理仿真音频;整体任务遵循四阶段流程,包括人工标注与基于人类表现的最终筛选。相比仅依赖图文回答导致准确率轻微下降的情况,STAR-Bench中模型在时间与空间推理上分别下降31.5%与35.2%,证明其聚焦于语言难以描述的感知线索。评估19个模型显示,人类与模型间存在显著差距,且呈现能力层级:闭源模型受细粒度感知限制,开放模型则在感知、知识与推理上全面滞后。本研究为发展具备更强物理世界理解力的未来模型提供了关键洞见与清晰路径。

原文摘要 · Abstract (English)

Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning. We formalize audio 4D intelligence that is defined as reasoning over sound dynamics in time and 3D space, and introduce STAR-Bench to measure it. STAR-Bench combines a Foundational Acoustic Perception setting (six attributes under absolute and relative regimes) with a Holistic Spatio-Temporal Reasoning setting that includes segment reordering for continuous and discrete processes and spatial tasks spanning static localization, multi-source relations, and dynamic trajectories. Our data curation pipeline uses two methods to ensure high-quality samples. For foundational tasks, we use procedurally synthesized and physics-simulated audio. For holistic data, we follow a four-stage process that includes human annotation and final selection based on human performance. Unlike prior benchmarks where caption-only answering reduces accuracy slightly, STAR-Bench induces far larger drops (-31.5\% temporal, -35.2\% spatial), evidencing its focus on linguistically hard-to-describe cues. Evaluating 19 models reveals substantial gaps compared with humans and a capability hierarchy: closed-source models are bottlenecked by fine-grained perception, while open-source models lag across perception, knowledge, and reasoning. Our STAR-Bench provides critical insights and a clear path forward for developing future models with a more robust understanding of the physical world.

音频理解时空推理多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。