构建化学奥赛新基准,用多智能体系统攻克复杂化学推理难题
ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025
- 设计多智能体框架模拟专家协作,分工处理问题分解、感知与推理
- 引入SVE机制分离视觉理解与化学推理,提升模型诊断能力
- 在国际化学奥赛2025数据集上达到93.6分,超人类金牌线
数学与物理的奥林匹克级基准对先进AI推理至关重要,但化学因其独特的多模态符号语言仍具挑战。我们基于2025年国际化学奥林匹克(IChO)推出新基准ChemO,包含两项关键创新:评估等价重构(AER),将需视觉输出的问题(如画分子结构)转化为可计算形式;结构化视觉增强(SVE),用于解耦模型的视觉感知能力与核心化学推理能力。为应对该基准,我们提出ChemLabs——一种分层多智能体系统,通过专用智能体实现问题分解、感知、推理与审计,模拟人类专家协作。实验表明,结合SVE与多智能体系统可显著提升性能。最优配置在ChemO上取得93.6分(满分100),超越预估的人类金牌门槛,确立自动化化学求解新标杆。
原文摘要 · Abstract (English)
Olympiad-level benchmarks in mathematics and physics are crucial testbeds for advanced AI reasoning, but chemistry, with its unique multimodal symbolic language, has remained an open challenge. We introduce ChemO, a new benchmark built from the International Chemistry Olympiad (IChO) 2025. ChemO features two key innovations for automated assessment: Assessment-Equivalent Reformulation (AER), which converts problems requiring visual outputs (e.g., drawing molecules) into computationally tractable formats, and Structured Visual Enhancement (SVE), a diagnostic mechanism to disentangle a model's visual perception capabilities from its core chemical reasoning. To tackle this benchmark, we propose ChemLabs, a hierarchical multi-agent framework that mimics human expert collaboration through specialized agents for problem decomposition, perception, reasoning, and auditing. Experiments on state-of-the-art multimodal models demonstrate that combining SVE with our multi-agent system yields dramatic performance gains. Our top configuration achieves a score of 93.6 out of 100, surpassing an estimated human gold medal threshold and establishing a new state-of-the-art in automated chemical problem-solving. ChemO Dataset: https://huggingface.co/datasets/IDEA-AI4SCI/ChemO
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。