arXiv:2608.30713cs.SD2026-08

通过自检机制提升长段落音频描述的准确性与忠实度。

Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning

论文配图:Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning
图 1 · 摘自论文原文
  • 将音频问答作为验证原语,贯穿生成全过程
  • 构建5万条长音频-长描述数据集,平均长度491.5词
  • 提出基于中间层曲率的微调方法,解决语义熵坍缩问题

长段落详细音频描述任务要求对细粒度音频内容进行密集且忠实于原文的描述,当前多模态语言模型仍无法有效完成。我们归因于两个结构性问题:一是数据贫乏,缺乏同时提供长片段、段落级描述和逐字转录忠实度的公开语料库;二是生成模式失败,表现为正确音频与打乱音频的多选题准确率差距达44.8至46.4个百分点。为此,我们提出Self-Check Captioning(SCC)统一框架,将音频接地的问答作为每个生命周期阶段的验证原语。SCC产出三个成果:构建了包含50,222个音频片段的长段落音频描述数据集(LACap-50k),平均描述长度为491.5词,并通过后置自动语音识别(ASR)审计验证;提出首个基于中间层证据权重的在策略监督微调方法(LC-SFT),源于对晚期层语义熵坍缩(SEC)的发现;设计SCC-Verifier,在推理时通过音频接地的自回答机制仲裁多个生成结果。在多个基准上,系统性能优于开源描述器,接近专有基线。我们公开发布LACap-50k以填补该研究领域的资源缺口。

原文摘要 · Abstract (English)

Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimodal language models. We attribute this failure to two structural problems. The first is data poverty, as no public corpus jointly provides long clips, paragraph captions, and verbatim-transcript fidelity. The second is generation-mode failure, evidenced by a 44.8 to 46.4 percentage-point gap between right-audio and shuffled-audio multiple-choice question (MCQ) accuracy. We address both within Self-Check Captioning (SCC), a unified framework that instantiates audio-grounded question answering as the verification primitive at every lifecycle stage. SCC yields three artifacts. Long-paragraph Audio Caption 50k (LACap-50k) is a 50,222-clip audio-visual corpus with 491.5-word captions and a post-hoc automatic speech recognition (ASR) audit. Layer-Curvature Supervised Fine-Tuning (LC-SFT) is the first on-policy supervised fine-tuning method to weight tokens by intermediate-layer evidence, motivated by our identification of Late-Layer Semantic-Entropy Collapse (SEC). SCC-Verifier arbitrates among caption rollouts via audio-grounded self-answering at inference. Across multiple benchmarks, our system attains state-of-the-art among open-source captioners and is competitive with proprietary baselines. We release LACap-50k to fill the resource gap for long-paragraph detailed audio captioning research.

音频描述自检机制数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。