用博弈论驱动多智能体框架,减少大模型科学推理幻觉
Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

- 融合贝叶斯与团队博弈,构建自适应多智能体闭环系统
- 生成36万条思维链和近20万问答对,7B模型幻觉降低79.46%
- 适用于化学分子设计等专业领域,提升科学推理可靠性
轻量级大语言模型在基于规则的科学领域应用受限,因其倾向于模仿语言模式而非遵循公理化推理,导致频繁幻觉。本文提出G-Frame,一种结合贝叶斯与团队博弈原则的自适应多智能体框架,实现高质量数据合成与模型训练的自动化闭环。通过结构化推理强制内化领域约束,我们构建了包含363,045条思维链和199,589个问答对的专业语料库。由此训练的7B模型OmniChem在自定义基准和ChemBench上表现媲美GPT-4o mini,同时相比基础架构幻觉率降低79.46%。进一步验证了OmniChem在分子设计与合成路径规划中的先进能力。本工作建立了一种可扩展的多智能体范式,有效克服大模型固有的推理缺陷,为加速专业化科学领域的知识发现提供了可行路径。
原文摘要 · Abstract (English)
The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic linguistic patterns rather than reproduce axiomatic reasoning, causing frequent hallucinations. Here, we show that G-Frame, an adaptive multi-agent framework integrating Bayesian and team game principles, establishes an automated closed-loop for high-quality data synthesis and model training. By forcing the internalization of domain constraints through structured reasoning, we synthesized a specialized corpus of 363,045 chains-of-thought and 199,589 question-answer pairs. The resulting 7B model OmniChem achieves performance parity with GPT 4o mini on custom benchmarks and ChemBench while exhibiting a 79.46% reduction in hallucinations relative to its base architecture. We further demonstrate the advanced capabilities of OmniChem in molecular design and synthesis planning. This work establishes a scalable paradigm utilizing adaptive multi-agents to overcome inherent reasoning deficiencies, offering a feasible pathway for accelerating knowledge discovery in specialized scientific fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。