arXiv:2602.15766cs.SD2026-02被引 3

TAC让音频描述精准定位时间,减少幻觉。

TAC: Timestamped Audio Captioning

  • 用合成数据训练,精准捕捉复杂声音场景中的事件时间点。
  • 在事件检测与密集描述任务中表现领先,幻觉率低。
  • 可连接大模型实现音视频理解推理,适合多模态研究者。

大型音频语言模型在复杂声学场景中难以区分重叠事件,导致时间描述不一致和频繁幻觉。我们提出时间标记音频描述模型TAC,可在不同细节层次和分辨率下生成具有时间定位的音频描述。TAC通过一个合成数据流水线构建来自真实音频源的动态复杂混合音频,从而在真实多音条件下实现稳健学习。在事件检测与密集描述任务中,TAC超越所有对比方法,具备低幻觉率和精确的时间定位能力。我们还引入TAC-V,一种音视频流水线,用于生成语义丰富的音视频描述。进一步证明,TAC和TAC-V可作为文本推理模型的“语义桥梁”:简单的TAC→LLM与TAC-V→LLM级联在音频(MMAU-Pro、MMSU、MMAR)和音视频(DailyOmni、VideoHolmes)理解与推理基准上均达到当前最优性能。

原文摘要 · Abstract (English)

Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection and dense captioning, TAC outperforms all competing methods, with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC and TAC-V serves as a "semantic bridge" for a text-only reasoner: a simple TAC$\rightarrow$LLM and TAC-V$\rightarrow$LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively.

音频描述时间定位多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。