构建首个结合全文上下文的科学图表质量评估基准,解决现有方法无法识别图文不一致的问题。
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

- 提出跨模态评估框架SFQ-Agent,融合图文与上下文信息进行多维度评分
- 在6308张科学图表上获取专家标注,验证了模型在5个维度上的高一致性
- 适用于期刊审稿、论文自动化质检及科研可视化工具开发
科学图像在呈现实验结论、阐述系统架构和支撑比较论证中至关重要。然而,现有图像质量评估方法主要针对自然照片或AI生成内容,难以直接应用于科学论文。现有研究仅限于视觉表面比对,无法验证图注一致性、引用相关性或视觉误导风险。为此,我们提出SciFigQual-Bench,一个基于全文本上下文的科学图表质量评估基准,涵盖清晰度、布局、图注匹配度、上下文相关性与误导风险五个维度。数据集覆盖2020至2025年顶级计算机科学会议论文,6,308张图像由多位领域专家独立评分并聚合为金标准标注。该数据集将每张图像与其图注、引用句及全文上下文绑定。为实现自动化评估,设计分阶段跨模态评估框架SFQ-Agent,通过多模态证据收集与融合实现可审计、精细化评分。在1200张测试图像上,配备GPT-5.6-Sol的SFQ-Agent(F3)取得最低平均绝对误差0.418和最高一致性率93.4%,显著优于直接评估与辅助视觉语言模型方案。
原文摘要 · Abstract (English)
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve auditable and refined scoring through the collection and fusion of modal evidence. Multiple mainstream large models were evaluated on the test subset eval1200, and SFQ-Agent (F3) equipped with GPT-5.6-Sol achieved the lowest overall average absolute error (0.418) and the highest consistency rate (93.4%), consistently outperforming both direct evaluation and auxiliary (Sidecar) visual language model evaluation schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。