用多个不同模型协作推理,显著提升AI答案的准确性和可靠性。
Collective Intelligence with Foundation Models

- 多个模型各自独立出方案,由评议员和聚合器协同优化。
- 异构模型组合使推理准确率提升至0.64,是单一模型的2.3倍。
- 适合需要高可信度、可解释性的科研与工业决策场景。
随着基础模型规模与多样性增长,将多个模型协调为协作推理系统成为实现更安全、更可靠AI的途径。本文提出多智能体框架:求解模型生成独立草稿,经评论员智能体结构化批判与修订,聚合器生成最终共识答案。评分模块对所有智能体进行语义、数值与流程评估。在涵盖微积分、物理、化学、生物、经济、优化、统计与数学的基准上进行消融实验,对比四种配置:(1)单个基线,(2)使用同一模型的同质框架,(3)多个相同模型的冗余同质求解器,(4)使用多样化专业模型的异构框架。结果表明,尽管框架结构与冗余采样带来小幅提升,但模型异质性是性能跃升的关键。异构配置在步骤级准确率上达0.64,较单模型提升2.3倍,且跨类别与难度水平方差降低。只有引入模型多样性,中间推理步骤的正确性才显著改善,证明异构智能体能互补纠错与优化推理,对可解释性与可审计性至关重要。本文探讨了架构原则、评估方法及对全球应用型AI的影响,表明异构多智能体协作可支撑科学与工业领域透明、可审计、高置信度决策。
原文摘要 · Abstract (English)
As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a path toward safer, more reliable AI. This chapter presents a multi-agent framework where solver models generate independent drafts, each undergoes structured critique and revision by a critic agent, and an aggregator agent synthesizes a final consensus solution. A scoring module provides semantic, numerical, and procedural evaluation across all agents. Through ablation studies on a benchmark spanning calculus, physics, chemistry, biology, economics, optimization, statistics, and mathematics, we isolate the contributions of framework architecture versus model diversity. We compare four configurations: (1) Individual Baseline, (2) Homogeneous Framework using one shared model, (3) Redundant Homogeneous Solvers using multiple instances of the same model, and (4) Heterogeneous Framework with diverse specialized models. Results show that while framework structure and redundant sampling yield modest gains, model heterogeneity is the critical factor driving substantial performance improvements. The heterogeneous configuration achieves superior step-wise accuracy (0.64 vs. 0.54 for individual models; 2.3x improvement over homogeneous configurations) with reduced variance across categories and difficulty levels. Step-wise reasoning quality (correctness of intermediate steps, not just final answers) improves dramatically only with model diversity, showing that heterogeneous agents provide complementary error detection and reasoning refinement essential for explainability and auditability. We discuss architectural principles, evaluation methodology, and implications for Global Applied AI, showing how heterogeneous multi-agent coordination supports transparent, auditable, high-confidence decision-making across scientific and industrial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。