提升视频字幕中对话描述的准确性,让AI更懂人物在说什么。
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
- 构建高质量数据集并分阶段优化对话描述能力
- 在新基准上超越Gemini系列模型的对话识别准确率
- 适合需要精准对话理解的多模态应用开发者
在音视频字幕生成中,准确描述对话对下游理解与生成任务至关重要。现有模型普遍难以生成忠实的对话内容。为此,我们提出DiaDem,一个能生成更精准对话描述且整体性能强的音视频字幕模型。首先,我们构建高质量SFT数据集;其次,采用难度分层的两阶段GRPO策略进一步优化对话描述。为系统评估对话描述能力,我们引入DiaDemBench,一个涵盖多样对话场景的综合评测基准,重点评估说话人识别准确率与话语转录保真度。在DiaDemBench上的大量实验表明,即使商业模型在对话感知字幕生成方面仍有显著提升空间。值得注意的是,DiaDem不仅在对话描述准确率上优于Gemini系列,同时在通用音视频字幕基准上也表现出竞争力,证明了其整体有效性。
原文摘要 · Abstract (English)
Accurate dialogue description in audiovisual video captioning is crucial for downstream understanding and generation tasks. However, existing models generally struggle to produce faithful dialogue descriptions within audiovisual captions. To mitigate this limitation, we propose DiaDem, a powerful audiovisual video captioning model capable of generating captions with more precise dialogue descriptions while maintaining strong overall performance. We first synthesize a high-quality dataset for SFT, then employ a difficulty-partitioned two-stage GRPO strategy to further enhance dialogue descriptions. To enable systematic evaluation of dialogue description capabilities, we introduce DiaDemBench, a comprehensive benchmark designed to evaluate models across diverse dialogue scenarios, emphasizing both speaker attribution accuracy and utterance transcription fidelity in audiovisual captions. Extensive experiments on DiaDemBench reveal even commercial models still exhibit substantial room for improvement in dialogue-aware captioning. Notably, DiaDem not only outperforms the Gemini series in dialogue description accuracy but also achieves competitive performance on general audiovisual captioning benchmarks, demonstrating its overall effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。