让大模型精准定位音频时间点,突破长音频理解瓶颈。
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
- 用时间标记和绝对时序编码增强时间感知能力
- 通过分段融合减少冗余,实现端到端长音频处理
- 构建新数据集与评估指标,专注细粒度时序任务
近期大型音频-语言模型在对话式问答中表现出色,但在时间定位(如时间音频定位)上表现不佳,且仅支持短音频感知,限制了其在细粒度任务中的应用。我们识别出三大制约因素:(i) 时间戳表示,(ii) 模型架构,(iii) 数据。为此,提出TimeAudio方法,通过引入独特的时间标记提升时序推理能力,并采用绝对时间感知编码,将声学特征与绝对时间信息显式对齐。为实现端到端长音频理解,设计分段级令牌融合模块,显著降低音频令牌冗余,提升信息提取效率。由于缺乏合适的数据集与评估指标,我们整合现有音频数据构建新数据集,专攻时序任务,并建立一系列评估指标。实验表明,TimeAudio在密集描述、时间定位、时间线语音摘要等多样化细粒度任务中表现优异,验证了其强大的时间定位与推理能力。
原文摘要 · Abstract (English)
Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perception, leading to constrained capabilities on fine-grained tasks. We identify three key aspects that limit their temporal localization and long audio understanding: (i) timestamp representation, (ii) architecture, and (iii) data. To address this, we introduce TimeAudio, a novel method that empowers LALMs to connect their understanding of audio content with precise temporal perception. Specifically, we incorporate unique temporal markers to improve time-sensitive reasoning and apply an absolute time-aware encoding that explicitly grounds the acoustic features with absolute time information. Moreover, to achieve end-to-end long audio understanding, we introduce a segment-level token merging module to substantially reduce audio token redundancy and enhance the efficiency of information extraction. Due to the lack of suitable datasets and evaluation metrics, we consolidate existing audio datasets into a new dataset focused on temporal tasks and establish a series of metrics to evaluate the fine-grained performance. Evaluations show strong performance across a variety of fine-grained tasks, such as dense captioning, temporal grounding, and timeline speech summarization, demonstrating TimeAudio's robust temporal localization and reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。