arXiv:2604.02934cs.CV2026-04被引 3

构建真实聚合物实验全流程评测基准,揭示大模型实践应用短板

PolyReal: A Benchmark for Real-World Polymer Science Workflows

  • 基于真实科研流程设计五类任务,覆盖从知识到实操的全链条
  • 大模型在机制推理中表现良好,但在安全分析与数据提取上显著失准
  • 适合关注科学AI落地、跨学科实证评估的研究者使用

多模态大语言模型在通用领域表现优异,但在复杂现实科学任务中表现欠佳。我们提出聚合物科学是一个理想且高风险的测试场景,因其横跨化学、物理、生物和工程等多学科,且数据形式多样。然而现有聚合物科学评测基准大多忽略真实实验流程,难以系统评估模型在完整科研生命周期中的表现。为此,我们构建了PolyReal——一个基于真实科研实践的多模态基准,涵盖五大核心能力:(1)基础知识应用;(2)实验室安全分析;(3)实验机理推理;(4)原始数据提取;(5)性能与应用探索。对主流MLLMs在PolyReal上的评估显示能力分布不均:模型在知识密集型任务(如实验机理推理)表现良好,但在实践导向任务(如安全分析、数据提取)显著下降,暴露出抽象科学知识与实际情境应用之间的严重脱节。这表明真实任务对当前大模型仍具挑战性。PolyReal填补了这一评估空白,为科学AI在真实工作流中的评估提供了实用基准。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) excel in general domains but struggle with complex, real-world science. We posit that polymer science, an interdisciplinary field spanning chemistry, physics, biology, and engineering, is an ideal high-stakes testbed due to its diverse multimodal data. Yet, existing benchmarks related to polymer science largely overlook real-world workflows, limiting their practical utility and failing to systematically evaluate MLLMs across the full, practice-grounded lifecycle of experimentation. We introduce PolyReal, a novel multimodal benchmark grounded in real-world scientific practices to evaluate MLLMs on the full lifecycle of polymer experimentation. It covers five critical capabilities: (1) foundational knowledge application; (2) lab safety analysis; (3) experiment mechanism reasoning; (4) raw data extraction; and (5) performance & application exploration. Our evaluation of leading MLLMs on PolyReal reveals a capability imbalance. While models perform well on knowledge-intensive reasoning (e.g., Experiment Mechanism Reasoning), they drop sharply on practice-based tasks (e.g., Lab Safety Analysis and Raw Data Extraction). This exposes a severe gap between abstract scientific knowledge and its practical, context-dependent application, showing that these real-world tasks remain challenging for MLLMs. Thus, PolyReal helps address this evaluation gap and provides a practical benchmark for assessing AI systems in real-world scientific workflows.

聚合物科学多模态评测科学AI大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。