提出科学图像质量评估新框架,兼顾内容正确性与认知清晰度。
SIQA: Toward Reliable Scientific Image Quality Assessment
- 从知识正确性和感知清晰度双维度定义科学图像质量
- 多模型测试显示评分一致但理解能力普遍不足
- 适合评估科研图像生成与审稿的学术研究者使用
科学图像不同于自然图像和AI生成图像,其核心是承载结构化领域知识而非仅呈现视觉场景。评估其质量需同时考量感知保真度、科学正确性与逻辑完整性。现有图像质量评估范式主要关注感知失真或图文对齐,隐含内容真实性的假设,在科学场景中失效——视觉上合理的图可能包含概念错误或推理不全。为此,我们提出科学图像质量评估(SIQA)框架,从知识(科学有效性与完整性)与感知(认知清晰度与学科规范性)两个互补维度建模质量。设计两种评估协议:SIQA-U(理解)通过多项选择任务衡量科学内容语义理解;SIQA-S(评分)评估与专家判断的一致性。构建包含专家标注基准与大规模训练集的SIQA挑战赛。在代表性多模态大模型上的实验表明,尽管模型在评分一致性上表现良好,但在理解任务中仍显著落后。微调虽能提升两项指标,但评分提升远超理解提升。结果表明,评分一致性不能可靠反映科学理解能力,强调科学图像评估需采用多维标准。
原文摘要 · Abstract (English)
Scientific images fundamentally differ from natural and AI-generated images in that they encode structured domain knowledge rather than merely depict visual scenes. Assessing their quality therefore requires evaluating not only perceptual fidelity but also scientific correctness and logical completeness. However, existing image quality assessment (IQA) paradigms primarily focus on perceptual distortions or image-text alignment, implicitly assuming that depicted content is factually valid. This assumption breaks down in scientific contexts, where visually plausible figures may still contain conceptual errors or incomplete reasoning. To address this gap, we introduce Scientific Image Quality Assessment (SIQA), a framework that models scientific image quality along two complementary dimensions: Knowledge (Scientific Validity and Scientific Completeness) and Perception (Cognitive Clarity and Disciplinary Conformity). To operationalize this formulation, we design two evaluation protocols: SIQA-U (Understanding), which measures semantic comprehension of scientific content through multiple-choice tasks, and SIQA-S (Scoring), which evaluates alignment with expert quality judgments. We further construct the SIQA Challenge, consisting of an expert-annotated benchmark and a large-scale training set. Experiments across representative multimodal large language models (MLLMs) reveal a consistent discrepancy between scoring alignment and scientific understanding. While models can achieve strong agreement with expert ratings under SIQA-S, their performance on SIQA-U remains substantially lower. Fine-tuning improves both metrics, yet gains in scoring consistently outpace improvements in understanding. These results suggest that rating consistency alone may not reliably reflect scientific comprehension, underscoring the necessity of multidimensional evaluation for scientific image quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。