用两阶段方法提升科学图像质量评估,兼顾视觉与科学信息。
SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment

- 分两阶段训练:先领域自适应预训练,再任务微调。
- 在科学图像质量数据集上取得92.21分(评分)和47.38分(理解)。
- 适合关注科学图像分析、评测系统研发的研究者。
科学图像对传达实验观察、量化证据和概念知识至关重要。与自然图像不同,其质量不仅取决于视觉清晰度,还依赖于科学信息量,评估难度高。本文提出SciQNet,一种用于科学图像质量评估的两阶段多模态适配框架。第一阶段在科学文档图像上进行领域自适应预训练,第二阶段通过联合评分与理解监督进行任务特定微调。评分监督结合指令微调与基于评分词概率的Huber损失,理解监督则采用多项选择式视觉问答。实验表明,在所测试的预训练数据比例中,使用40%的分层子集表现最佳,提示预训练数据的相关性可能比规模更重要。最终模型在SIQA-S上得分为92.21,SIQA-U为47.38,综合得分为69.80。该工作是针对ICME 2026科学图像质量评估挑战赛的解决方案,在评分赛道中排名第二。
原文摘要 · Abstract (English)
Scientific images are essential for communicating experimental observations, quantitative evidence and conceptual knowledge. Unlike natural images, their quality depends on both visual clarity and scientific informativeness, making assessment challenging. In this work, we present SciQNet, a two-stage multimodal adaptation framework for scientific image quality assessment. The first stage performs domain-adaptive pretraining on scientific document images and the second stage conducts task-specific fine-tuning with joint scoring and understanding supervision. For scoring-oriented supervision, we combine instruction tuning with a Huber loss derived from rating-word logits, while understanding-oriented supervision is formulated as multiple-choice visual question answering. Experiments show that using a 40% stratified subset of the domain-adaptive data gives the best performance among the evaluated pretraining fractions, suggesting that pretraining-data relevance may be as important as pretraining-data scale. The final model achieves an SIQA-S score of 92.21, an SIQA-U score of 47.38 and a combined score of 69.80. This work presents our solution to the ICME 2026 Scientific Image Quality Assessment Challenge, which ranked 2nd in the scoring track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。