arXiv:2608.24160cs.AI2026-08

提出新基准D3-Omni,揭示多模态评判模型的隐藏盲区。

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

  • 构建平衡解耦的评测集,分离53个独立判别维度
  • 发现强模型在识别违规项上远弱于满足项
  • 适合评估和改进多模态生成模型的判别能力

多模态理解模型常被用作跨文本到图像(T2I)、文本到视频(T2V)和文本到语音(TTS)生成的统一评判者(OmniJudge)。然而,现有评测数据和训练集偏向正样本,且混杂不同失败模式,导致模型得分高却不真正识别缺陷。为此,我们提出D3-Omni基准,涵盖53个正交二元维度(T2I:17, T2V:22, TTS:14),共10,671个样本(T2I:3,526, T2V:1,998, TTS:5,147)。通过固定已验证的全正样本种子,采用受控提示重写与原子化、维度隔离扰动生成负样本,避免维度间信息泄露。D3设计具双重平衡性:缓解负样本稀缺与维度标签失衡;解耦性确保每类错误可归因单一能力;动态构造机制引导向生成模型改进后仍不足的标签区域。该评测集实现各维度近乎1:1的准确率与总分层级的均匀分布。在这一平衡视角下,即使强模型也普遍在模态相关维度表现不佳,更可靠地确认满足条件而非检测违反条件,且将名义上不同的属性视为单一决策,表明整体准确率可能掩盖系统性盲点,而平衡解耦的评测视角能有效暴露并助力修复这些缺陷。

原文摘要 · Abstract (English)

Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve.The suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.

多模态评测模型诊断解耦评估生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。