arXiv:2608.21357cs.AI2026-08

构建生命科学视觉分析基准,测试AI对实验图像的理解能力。

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

论文配图:VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
图 1 · 摘自论文原文
  • 设计161个真实科研流程中的图像理解任务
  • 现役视觉语言模型在科学图像上表现不佳
  • 适合评估AI在生物技术领域的实用能力

在专业生命科学工作中,科学家常通过凝胶电泳图、显微图像、质粒图谱、流式细胞图、分子结构图等视觉图像辅助研究决策。我们提出VIALS,一个包含161个此类图像理解任务的视觉问答基准,覆盖生物技术产业中实验全流程涉及的图像类型(而非出版物或教科书中的美化图像)。尽管前沿视觉-语言模型能流畅描述自然图像,但我们在科学图像上发现其解读准确率低下,反映出领域知识与特定视觉推理能力的不足。相比之下,具备相关领域经验的科学家可轻松完成这些任务。若AI无法有效理解此类图像,将难以在以图像为核心的研究决策流程中发挥作用。

原文摘要 · Abstract (English)

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.

图像理解生命科学视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。