首个评估多模态指令遵循能力的基准,发现格式越复杂模型越难理解。
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

- 构建多模态指令遵循评测框架,分格式与内容双维度评估。
- 在1920个样本上测试,发现格式复杂度与内容理解能力呈负相关。
- 开源54K指令微调数据集,助力模型更好遵循复杂指令。
尽管多模态大语言模型(OLLMs)在联合处理音视频流方面表现优异,但其严格遵循复杂多维用户指令的能力仍缺乏系统研究。现有基准多聚焦于整体视频理解或纯文本指令跟随,未能捕捉模态间与用户约束的复杂互动。为此,我们提出OmniCap-IF,首个专为多模态字幕生成设计的指令遵循评测基准。该基准采用系统化框架,从格式正确性与内容正确性两个维度评估字幕质量,涵盖纯视觉、纯音频及音视频融合模态下的50种不同约束类型,并引入时间定位(Temporal Grounding)以评估时空精度。对1,920个高质量样本的广泛测评揭示显著性能差异。分析发现存在关键的“格式-内容权衡”现象:格式复杂度越高,模型的多模态推理能力越弱。为推动领域发展,我们构建了包含54K样本的指令微调数据集OmniCap-IF-54K,提出OmniCaptioner-IF,在复杂指令遵循与通用多模态字幕生成上均取得显著提升。
原文摘要 · Abstract (English)
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex, multi-faceted user instructions remains largely unexplored. Existing benchmarks primarily focus on holistic video understanding or text-only instruction following, failing to capture the intricate interplay between modalities and user constraints. To bridge this gap, we introduce OmniCap-IF, the first comprehensive benchmark specifically designed to evaluate instruction-following capabilities in omni-modal captioning. OmniCap-IF incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our benchmark encompasses 50 distinct constraint types across pure visual, pure audio, and audio-visual modalities, while integrating Temporal Grounding to assess spatio-temporal precision. Extensive evaluations of prominent models on 1,920 high-quality samples reveal significant performance disparities. Furthermore, our analysis uncovers a critical "format-content tradeoff", demonstrating that increasing formatting complexity directly degrades models' omni-modal reasoning abilities. Finally, to advance the field, we curate a 54K instruction-tuning dataset, OmniCap-IF-54K and present OmniCaptioner-IF, which achieves notable improvements in both complex instruction adherence and general omni-modal captioning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。