用视觉语言模型提升科学发现的自主性,让系统自动纠错并实时调整探索方向。
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
- 用视觉语言模型作为裁判,动态评估图表并指导多智能体自我修正
- 在10项数据驱动任务中,准确率提升至0.7-0.8,远超代码或图文基线
- 生成可审计的推理过程,适合需要透明性和自适应能力的研究场景
我们证明,由视觉语言模型(VLMs)引导的多智能体系统能提升端到端的自主科学发现能力。通过将图表视为可验证的检查点,VLM作为评判者,依据动态生成的领域特定评分标准评估图像,使智能体能够自我纠正错误,并实时引导探索性数据分析。在宇宙学和天体化学的案例研究中,系统成功从错误推理路径中恢复,并在无须人工干预的情况下适应新数据集。在一项包含10个任务的数据驱动发现基准测试中,增强VLM的系统取得0.7-0.8的通过率(pass@1),显著优于仅使用代码的基线(0.2-0.3)和代码+文本基线(0.4-0.5),同时提供可审计的推理轨迹,提升了可解释性。代码已公开:https://github.com/CMBAgents/cmbagent
原文摘要 · Abstract (English)
We show that multi-agent systems guided by vision-language models (VLMs) improve end-to-end autonomous scientific discovery. By treating plots as verifiable checkpoints, a VLM-as-a-judge evaluates figures against dynamically generated domain-specific rubrics, enabling agents to correct their own errors and steer exploratory data analysis in real-time. Case studies in cosmology and astrochemistry demonstrate recovery from faulty reasoning paths and adaptation to new datasets without human intervention. On a 10-task benchmark for data-driven discovery, VLM-augmented systems achieve pass at 1 scores of 0.7-0.8, compared to 0.2-0.3 for code-only and 0.4-0.5 for code-and-text baselines, while also providing auditable reasoning traces that improve interpretability. Code available here: https://github.com/CMBAgents/cmbagent
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。