构建了多模态细节感知的全流程体系,提升音视频描述精度与真实性。
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
- 用智能工具自动生成高细节、低幻觉的多模态数据
- 音频/音视频细粒度描述模型性能超开源方案,媲美顶级闭源模型
- 设计新评测基准Omni-Cloze,稳定评估多模态细节生成能力
细粒度多模态感知对提升人机交互至关重要。随着音视频技术发展,可并行处理音频与视频信号的通用语言模型(OLMs)成为实现更丰富理解与推理的有前景范式,但其捕捉与描述细粒度细节的能力仍受限。本文从数据管道、模型与评测三方面系统研究多模态细节感知。发现当前OLMs存在细节与幻觉的内在‘共生长’问题。为此提出Omni-Detective,一种融合工具调用的智能数据生成管道,可自主生成高细节且低幻觉的多模态数据。基于该数据训练出两个描述模型:Audio-Captioner(纯音频)与Omni-Captioner(音视频)。在级联评估协议下,Audio-Captioner在MMAU和MMAR指标上超越所有开源模型,表现接近Gemini 2.5 Pro。在现有细粒度描述基准上,Omni-Captioner在VDC上达到新SOTA,且在video-SALMONN 2测试集上实现最佳细节与幻觉平衡。鉴于缺乏专用评测基准,本文设计Omni-Cloze——一种新型填空式评测方法,用于稳定高效评估音频、视觉及音视频细粒度描述,实验验证其有效性。
原文摘要 · Abstract (English)
Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer understanding and reasoning. However, their capacity to capture and describe fine-grained details remains limited explored. In this work, we present a systematic and comprehensive investigation of omni detailed perception from the perspectives of the data pipeline, models, and benchmark. We first identify an inherent "co-growth" between detail and hallucination in current OLMs. To address this, we propose Omni-Detective, an agentic data generation pipeline integrating tool-calling, to autonomously produce highly detailed yet minimally hallucinatory multimodal data. Based on the data generated with Omni-Detective, we train two captioning models: Audio-Captioner for audio-only detailed perception, and Omni-Captioner for audio-visual detailed perception. Under the cascade evaluation protocol, Audio-Captioner achieves the best performance on MMAU and MMAR among all open-source models, surpassing Gemini 2.5 Flash and delivering performance comparable to Gemini 2.5 Pro. On existing detailed captioning benchmarks, Omni-Captioner sets a new state-of-the-art on VDC and achieves the best trade-off between detail and hallucination on the video-SALMONN 2 testset. Given the absence of a dedicated benchmark for omni detailed perception, we design Omni-Cloze, a novel cloze-style evaluation for detailed audio, visual, and audio-visual captioning that ensures stable, efficient, and reliable assessment. Experimental results and analysis demonstrate the effectiveness of Omni-Detective in generating high-quality detailed captions, as well as the superiority of Omni-Cloze in evaluating such detailed captions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。