arXiv:2606.03401cs.CV2026-06被引 1

提出科学图像评估新框架,解决AI生成图像的隐性错误检测与修复难题。

Towards Characterizing Scientific Image Utility and Upgradability

论文配图:Towards Characterizing Scientific Image Utility and Upgradability
图 1 · 摘自论文原文
  • 构建误差类型分类体系,涵盖细节失真、信息缺失等四类问题
  • 多模态模型在错误识别与修复上准确率不足40%,存在严重能力缺口
  • 适合科研人员、期刊审稿人及AI内容安全研究者使用

科学图像在研究传播中是关键证据,但面临AI生成内容带来的细微却严重的错误威胁。现有评估方法效果有限:视觉质量指标与科学有效性相关性弱,语言模型缺乏领域验证能力。为此,我们提出科学图像实用性和可升级性评估(SIU²A)框架,包含两个互补维度:实用性涵盖错误检测(识别科学性偏差)和修复可行性(判断能否可靠修复);可升级性衡量修复后的质量。将科学图像损坏分为四类:细节失真、不完整、虚假内容、实体混淆。基于此构建了带专家标注的SIU²A-Benchmark数据集。框架采用双阶段评估:第一阶段评估错误检测与修复指令生成能力;第二阶段检验修复是否忠实还原科学真实性且不破坏原有正确信息。实验表明,当前多模态系统在科学错误评估和忠实修复方面均表现不佳,准确率低于40%,暴露出视觉感知与科学可用性之间的根本差距。

原文摘要 · Abstract (English)

Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but consequential errors. Existing evaluation paradigms prove inadequate: perceptual quality metrics poorly correlate with scientific validity, while language models lack domain-specific verification capabilities. To address this gap, we propose the \textbf{S}cientific \textbf{I}mage \textbf{U}tility and \textbf{U}pgradability \textbf{A}ssessment (\textbf{SIU$^2$A}) framework, which introduces two complementary dimensions for scientific image evaluation. \textbf{Utility} encompasses \textit{error detection} (identifying scientific inaccuracies) and \textit{correction feasibility} (assessing whether errors can be reliably repaired). \textbf{Upgradability} measures the quality of correction. We categorize scientific image corruption into four fundamental types: Detail Distortion, Incompleteness, False Content, and Entity Confusion. Based on this taxonomy, we construct SIU$^2$A-Benchmark, a dataset with expert annotations for error identification and repair. The framework implements a two-stage evaluation protocol: the \textit{Utility} stage evaluates error detection capability and repair instruction generation, while the \textit{Upgradability} stage assesses whether corrections faithfully restore scientific validity without compromising existing accurate information. Experiments reveal that current multimodal systems exhibit significant limitations in both scientific error assessment and faithful correction, exposing a fundamental gap between visual perception and scientific usability.

图像评估AI可信度科学传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。