arXiv:2604.22245eess.AS2026-04被引 3

解决长音频时间感知难题,提升模型对长时间音频的精准理解能力。

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

论文配图:Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding
图 1 · 摘自论文原文
  • 构建全局-局部渐进式推理框架,结合时间语义上下文与分步推理。
  • 在30分钟长音频上超越现有模型,时间对齐准确率显著提升。
  • 适合研究长音频理解、多模态时序建模的学者和工程师。

尽管大音频语言模型(LALMs)在短音频上表现良好,但在长音频输入上性能下降,尤其在时间感知任务中,随着音频长度增加,时间对齐准确性持续降低。我们将其归因于缺乏针对长音频时间感知的数据、基准和建模方法。为此,我们构建了包含1.2千小时真实场景音频的LAT-Chronicle数据集,以及支持长达30分钟音频的首个人工验证基准LAT-Bench,涵盖密集音频描述、时间音频定位和目标音频描述三项核心任务。基于这些资源,我们提出LAT-Audio,将时间感知建模为从全局到局部的渐进式推理范式:首先构建对齐的时间-语义上下文,再通过引入思考-听觉思维链(TWA-CoT),利用工具迭代融合局部音频信息进行推理。实验表明,LAT-Audio在长音频时间感知任务中优于现有模型,并增强了对输入长度变化的鲁棒性。相关数据集、基准与模型已开源,以推动后续研究。

原文摘要 · Abstract (English)

While Large Audio Language Models (LALMs) achieve strong performance on short audio, they degrade on long-form inputs. This degradation is more severe in temporal awareness tasks, where temporal alignment becomes increasingly inaccurate as audio duration grows. We attribute these limitations to the lack of data, benchmarks, and modeling approaches tailored for long-form temporal awareness. To bridge this gap, we first construct LAT-Chronicle, a 1.2k hour long-form audio dataset with temporal annotations across real-world scenarios. We further develop LAT-Bench, the first human-verified benchmark supporting audio up to 30 minutes while covering three core tasks: Dense Audio Caption, Temporal Audio Grounding, and Targeted Audio Caption. Leveraging these resources, we propose LAT-Audio, formulating temporal awareness as a progressive global-to-local reasoning paradigm. A global timeline is first constructed as an aligned temporal-semantic context,and the Think-With-Audio Chain-of-Thought (TWA-CoT) is then introduced to perform iterative reasoning by incorporating local audio information via tool use. Experiments show that LAT-Audio surpasses existing models on long-form audio temporal awareness tasks and improves robustness to input duration. We release the dataset, benchmark, and model to facilitate future research at https://github.com/alanshaoTT/LAT-Audio-Repo.

长音频时间感知推理框架多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。