首个基于真实审稿意见的多模态论文不一致检测基准,测试模型跨模态推理能力。
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- 从353篇论文中提取384个真实审稿人标记的多模态不一致问题
- 21个主流大模型平均表现仅27.8%-53.9%,说明科学推理能力严重不足
- 采用结构化答案格式避免答题套路,更真实评估理解能力
大型多模态模型(LMMs)在科研中的应用日益广泛,但其是否能可靠理解并推理论文中的多模态复杂性仍不明确。核心挑战在于检测和解决文本、图表、表格与公式间的不一致,这类问题往往细微、领域特定,最终损害清晰度、可复现性和可信度。现有基准要么孤立单模态,要么依赖合成错误,无法反映真实复杂性。我们提出PRISMM-Bench(基于审稿人反馈的多模态不一致数据集),通过多阶段流程(审稿挖掘、LLM辅助筛选、人工验证),从353篇论文中收集384个真实不一致。基于此,设计三项任务:不一致识别、修复建议和配对匹配,评估模型跨模态检测、修正与推理能力。为解决多选题中‘选择捷径’问题(模型依赖答案模式而非真正理解),引入结构化JSON答案表示,减少语言风格偏差。我们评测了21个领先LMMs,包括开源模型(GLM-4.5V 106B、InternVL3 78B)和专有模型(Gemini 2.5 Pro、GPT-5高推理版)。结果表明性能极低(27.8%-53.9%),凸显多模态科学推理的严峻挑战,推动可信科研助手的发展。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) are increasingly applied to scientific research, yet it remains unclear whether they can reliably understand and reason over the multimodal complexity of papers. A central challenge lies in detecting and resolving inconsistencies across text, figures, tables, and equations, issues that are often subtle, domain-specific, and ultimately undermine clarity, reproducibility, and trust. Existing benchmarks overlook this issue, either isolating single modalities or relying on synthetic errors that fail to capture real-world complexity. We introduce PRISMM-Bench (Peer-Review-sourced Inconsistency Set for Multimodal Models), the first benchmark grounded in real reviewer-flagged inconsistencies in scientific papers. Through a multi-stage pipeline of review mining, LLM-assisted filtering and human verification, we curate 384 inconsistencies from 353 papers. Based on this set, we design three tasks, namely inconsistency identification, remedy and pair matching, which assess a model's capacity to detect, correct, and reason over inconsistencies across different modalities. Furthermore, to address the notorious problem of choice-only shortcuts in multiple-choice evaluation, where models exploit answer patterns without truly understanding the question, we further introduce structured JSON-based answer representations that minimize linguistic biases by reducing reliance on superficial stylistic cues. We benchmark 21 leading LMMs, including large open-weight models (GLM-4.5V 106B, InternVL3 78B) and proprietary models (Gemini 2.5 Pro, GPT-5 with high reasoning). Results reveal strikingly low performance (27.8-53.9\%), underscoring the challenge of multimodal scientific reasoning and motivating progress towards trustworthy scientific assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。