arXiv:2607.10387eess.AScs.CL2026-07中稿 · Interspeech 2026

让大模型能精准定位长达2小时音频中的时间点。

GigaChat Audio: Time-aware Large Audio Language Model

论文配图:GigaChat Audio: Time-aware Large Audio Language Model
图 1 · 摘自论文原文
  • 用周期性时间标记混合连续音频令牌,实现时间感知
  • 在短长音频上均达到高时间定位准确率,支持带时间锚点的摘要
  • 适合需要精确时间定位的语音分析、智能剪辑等场景

长音频中的时间定位对音频条件大模型仍是挑战。本文提出一种时间感知音频大模型,可在长达120分钟的输入中回答带明确时间戳的问题。方法通过级联流水线生成大规模合成监督数据,将周期性时间标记与连续音频令牌交错处理。模型在短时和长时基准测试中均表现出强时间定位准确性,支持带时间锚点的片段描述与摘要生成。大量消融实验分析了时间表示方式、标记频率、分词策略及时长混合设计对准确率与计算成本的影响。模型权重与数据集已公开,可访问 https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B 支持后续研究。

原文摘要 · Abstract (English)

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.

音频理解时间定位大模型语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。