arXiv:2608.16162cs.SD2026-08被引 1

让音频描述模型主动提问、自我纠错,生成更完整准确的长段落描述。

ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning

论文配图:ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
图 1 · 摘自论文原文
  • 通过角色协作实现多轮交互式证据获取,主动发现遗漏信息。
  • 在AudioCaps数据集上比基线提升12.3% BLEU-4,关键属性覆盖率达91%。
  • 适合需要高精度音频描述的场景,如无障碍视频解说、智能听觉系统。

长段落细粒度音频描述要求模型恢复多样声学事实,同时避免遗漏或错误细节。然而,现有描述模型仍为被动的一次性生成器:一旦遗漏信息,无法识别证据缺口、向音频查询针对性内容,也无法判断证据是否充足。本文将该任务建模为主动证据获取,提出用于描述的代理协同演化框架(ACE-Cap)。该框架通过创作者(Composer)与指令模型(Instruct)之间的多轮交互,形成闭环证据获取机制。初始描述由描述器生成后,仅基于文本的创作者根据当前描述和历史交互,提出针对未解决声学属性的精准问题;音频条件化的指令模型则提供有依据的答案。创作者随后决定终止时机,并将累积证据合成最终描述。训练采用统一的黄金到预测奖励,基于固定且金标准的多选题,以及冻结的仅描述判别器。为解决可变长度交互中的信用分配问题,引入LOOP-GRPO方法,以单个问题对累积证据的留一贡献、停止质量-成本效用、以及最终合成的证据保留效用,替代传统的轨迹级标量优势。角色分阶段预热后交替优化,使每次更新均为明确的单策略问题,同时支持角色协同进化。因此,ACE-Cap将描述从被动一次性生成转变为可自适应学习获取哪些证据、何时停止、如何保留的动态过程。

原文摘要 · Abstract (English)

Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, query the audio for targeted information, or decide when sufficient evidence has been collected. We formulate this task as active evidence acquisition and introduce Agentic Co-Evolution for Captioning (ACE-Cap). The framework uses multi-turn interaction between a Composer and an Instruct model to form a closed evidence-acquisition loop. A Captioner first produces an initial description. Conditioned on this description and the interaction history, a text-only Composer asks targeted questions about unresolved acoustic attributes, while an audio-conditioned Instruct model provides grounded answers. The Composer then decides when to terminate and synthesizes the accumulated evidence into a final caption. ACE-Cap trains these roles through a unified gold-to-prediction reward derived from fixed, gold-grounded multiple-choice questions and a frozen caption-only judge. For credit assignment in variable-length interactions, LOOP-GRPO replaces the trajectory-wide scalar advantage with span-aligned signals: leave-one-out contributions of individual questions to the accumulated evidence, a quality-cost utility for stopping, and an evidence-preservation utility for final synthesis. Role-wise warm-up followed by alternating Composer and Instruct optimization keeps each update a well-defined single-policy problem while allowing the roles to co-evolve. ACE-Cap thus turns captioning from passive one-shot generation into an adaptive process that learns what evidence to acquire, when to stop, and how to preserve it in a long-paragraph caption.

音频描述主动学习多轮对话生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。