用严格视觉事实核查提升多模态模型评估真实可靠性
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

- 将评估从整体语义匹配转为原子级事实审计,设计可细化的评分准则
- 发现模型在密集场景中常因多重约束失败而表现脆弱,存在8%感知差距
- 引入门控打分机制,更贴近人类感知,适合高精度生成任务评估
我们提出PerceptionRubrics,一种基于评分表的多模态评估框架,弥补基准分数饱和与真实世界脆弱性之间的差距。该框架将评估从整体语义匹配转向严格的原子级审计,配对1,038张信息密集图像与超过10,000条实例特定评分准则。这些准则源自通过新型环形同行评审共识流程构建的黄金标注,并提炼为双流体系:必须正确(核心事实)与易出错(细粒度细节)。关键的是,PerceptionRubrics采用门控打分机制:不同于线性平均,关键视觉事实失败即触发尖锐二值惩罚。大量评估揭示三大洞见:(1) 可靠性缺口:模型常正确识别零散元素,却在严格联结约束下失败,暴露密集领域中的脆弱性;(2) 开源-专有分层:与推理趋势相反,我们发现开源与专有模型间持续存在8%的感知差距;(3) 人类对齐严谨性:门控指标显著优于传统基准,验证严格感知保真度是可靠生成的前提。
原文摘要 · Abstract (English)
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from holistic semantic matching to rigorous atomic auditing, PerceptionRubrics pairs 1,038 information-dense images with over 10,000 instance-specific rubrics. These criteria are derived from golden captions constructed via a novel Circular Peer-Review consensus pipeline and then distilled into a dual-stream system of Must-Right (essential facts) and Easy-Wrong (fine-grained details) rubrics. Crucially, PerceptionRubrics implements a Gated Scoring mechanism: unlike linear averages, failure on mandatory visual facts triggers sharp binary penalties. Extensive evaluation yields critical insights: (1) The Reliability Gap: models often verify fragmented elements correctly yet fail strict conjunctive constraints, exposing brittleness in dense domains; (2) Open-Closed Stratification: contrary to reasoning trends, we reveal a persistent 8% perception deficit between open-source and proprietary frontiers; and (3) Human-Aligned Rigor: our gated metrics substantially out-align conventional benchmarks, validating that strict perceptual fidelity is the prerequisite for reliable generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。