arXiv:2512.16921cs.CVcs.AI2025-12被引 1

用自动审计发现大模型能力短板并修复,效果显著。

Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification

  • 训练一个智能审计员,通过生成挑战性问题和反事实图像来暴露模型差异。
  • 在顶尖模型上发现20多种故障类型,微调后16个基准测试全提升。
  • 无需人工标注,30亿参数模型可超越280亿参数对手,适合模型优化者。

传统多模态大模型评估方法缺乏可解释性,难以全面揭示模型间的能力差距。为此,我们提出AuditDM——一种自动化框架,通过审计模型间的差异来主动发现并修复其失败模式。AuditDM利用强化学习微调多模态大模型作为审计员,使其生成能最大化目标模型分歧的挑战性问题与反事实图像。训练完成后,该审计员能揭示多样且可解释的典型样本,暴露出模型弱点,并作为无标注数据用于模型修正。应用于Gemma-3和PaliGemma-2等领先模型时,AuditDM发现了超过20种不同的失败类型。基于这些发现进行微调,显著提升了所有模型在16个基准上的表现,并使一个30亿参数模型超越280亿参数的同类模型。结果表明,当数据规模增长进入边际效应递减阶段,针对性模型审计成为诊断与改进的有效路径。

原文摘要 · Abstract (English)

Conventional evaluation methods for multimodal LLMs (MLLMs) lack interpretability and are often insufficient to fully disclose significant capability gaps across models. To address this, we introduce AuditDM, an automated framework that actively discovers and rectifies MLLM failure modes by auditing their divergence. AuditDM fine-tunes an MLLM as an auditor via reinforcement learning to generate challenging questions and counterfactual images that maximize disagreement among target models. Once trained, the auditor uncovers diverse, interpretable exemplars that reveal model weaknesses and serve as annotation-free data for rectification. When applied to SoTA models like Gemma-3 and PaliGemma-2, AuditDM discovers more than 20 distinct failure types. Fine-tuning on these discoveries consistently improves all models across 16 benchmarks, and enables a 3B model to surpass its 28B counterpart. Our results suggest that as data scaling hits diminishing returns, targeted model auditing offers an effective path to model diagnosis and improvement.

模型审计能力诊断大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。