让多个AI模型像专家辩论一样协作,用新协议发现隐藏认知盲点。
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis
- 给AI分配角色人格,分离能力与推理方式,让分歧变有价值。
- 低成本模型表现媲美高价模型,且能发现训练数据看不到的盲点。
- 适合需要客观分析、避免偏见的研究或政策制定场景。
我们提出Consilium协议,一种基于拜占庭容错的多模型协同推理架构,将模型间分歧视为知识信号而非错误。该协议为语言模型赋予工程化认知人格,实现模型能力与其推理方式的解耦,并引入源自量化金融的样本内/样本外验证框架,以区分训练数据共识与实证结论。在涵盖10个领域32个主题的1,478次辩论中,我们发现:(1) 认知人格决定知识行为,成本仅0.0002美元/批的边缘推理模型输出效果可媲美成本10.69美元的前沿模型;(2) RLHF对齐训练产生可测量的领域特异性认知盲区——争议性议题的对抗性挑战比科学共识议题低12.3个百分点,而人工智能安全议题呈现不对称偏见(Δ=11.6%),模型更激烈质疑AI危险性说法;(3) 协议自身无方向偏差(移民Δ=2.3%,可再生能源Δ=1.2%);(4) 样本外证据检索成功验证239条主张(100%召回),并发现167项训练数据无法察觉的认知盲点。多次运行间重现性标准差平均±2.2%。完整实验总成本217美元。协议规范已开源,采用MIT许可,支持独立验证。
原文摘要 · Abstract (English)
We present the Consilium Protocol, a Byzantine Fault Tolerance-derived architecture for structured multi-model AI deliberation that treats inter-model disagreement as epistemic signal rather than error. The protocol assigns engineered cognitive personas to language models -- separating what a model is from how it reasons -- and introduces an In-Sample/Out-of-Sample validation framework adapted from quantitative finance to distinguish training-data consensus from empirically grounded conclusions. Across 1,478 deliberation sessions spanning 32 topics in 10 domain categories, we demonstrate that (1) the cognitive persona, not the underlying model, determines epistemic behavior: free edge-inference models costing 0.0002 USD per batch produced comparable analytical output to frontier models costing 10.69 USD; (2) RLHF alignment training creates measurable, domain-specific epistemic blind spots -- contested policy topics exhibit 12.3 percentage points less adversarial challenge than settled science topics, and AI safety topics show asymmetric bias ($Δ$=11.6%) where models challenge claims that AI is dangerous far more vigorously than claims that AI risk is overstated; (3) the protocol exhibits no directional bias of its own (immigration $Δ$=2.3%, renewables $Δ$=1.2%); and (4) out-of-sample evidence retrieval validated 239 claims with 100% evidence retrieval and surfaced 167 blind-spot discoveries invisible to training-data deliberation. Run-to-run reproducibility across randomized model$\times$persona assignments averages $\pm$2.2% standard deviation. Total cost for the complete battery including all overhead: 217 USD. We release the protocol specification under MIT license to enable independent verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。