arXiv:2602.08979cs.SDcs.CL2026-02ACL被引 2

提出纯音频模型AudioSeg,显著提升长音频分段效果。

Beyond Transcripts: A Renewed Perspective on Audio Chaptering

  • 用音频特征替代文本,设计纯音频分段模型AudioSeg
  • 停顿信息带来最大性能提升,文本错误影响明显
  • 适合处理无字幕长音频,尤其对语音内容研究者有用

音频分段任务将长音频划分为语义连贯的片段,在播客、讲座和视频中日益重要。现有研究多依赖文本,缺乏对音频信息利用、语音识别错误处理及无字幕评估方法的探讨。本文通过三项贡献填补空白:(1)系统比较基于文本+声学特征的模型、全新的纯音频架构AudioSeg(基于学习的音频表示)以及多模态大模型;(2)实证分析影响性能的关键因素,包括转录质量、声学特征、音频时长与说话人构成;(3)建立形式化评估协议,对比依赖转录的文本空间方法与不依赖转录的时间空间方法。在YTSeg数据集上的实验表明,AudioSeg显著优于文本基方法,停顿提供最大音频增益,多模态大模型受上下文长度限制且指令遵循较弱,但在短音频上仍具潜力。

原文摘要 · Abstract (English)

Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation. We address these gaps through three contributions: (1) a systematic comparison between text-based models with acoustic features, a novel audio-only architecture (AudioSeg) operating on learned audio representations, and multimodal LLMs; (2) empirical analysis of factors affecting performance, including transcript quality, acoustic features, duration, and speaker composition; and (3) formalized evaluation protocols contrasting transcript-dependent text-space protocols with transcript-invariant time-space protocols. Our experiments on YTSeg reveal that AudioSeg substantially outperforms text-based approaches, pauses provide the largest acoustic gains, and MLLMs remain limited by context length and weak instruction following, yet MLLMs are promising on shorter audio.

音频分段纯音频多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。