arXiv:2509.00091cs.AIcs.CL2025-09

用本地开源模型组辩论提升AI对齐能力,效果优于单模型。

Ensemble Debates with Local Large Language Models for AI Alignment

  • 构建本地开源模型组进行辩论,通过多视角推理增强对齐性。
  • 辩论组在推理深度和论点质量上分别提升19.4%和34.1%。
  • 适合关注可复现性与开放评估的AI对齐研究者使用。

随着大语言模型在高风险决策中作用日益突出,其与人类价值观的对齐至关重要。依赖专有API限制了可复现性和广泛参与。我们研究了本地开源模型组辩论是否能提升对齐导向的推理能力。在涵盖15种场景、5种组构型的150场辩论中,组辩论在7分制评分中表现更优(整体得分3.48 vs. 3.13),推理深度提升19.4%,论点质量提升34.1%。改进最显著的是真实性(+1.25分)和人类增强性(+0.80分)。我们公开代码、提示模板和辩论数据集,为基于组辩论的对齐评估提供可访问且可复现的基础。

原文摘要 · Abstract (English)

As large language models (LLMs) take on greater roles in high-stakes decisions, alignment with human values is essential. Reliance on proprietary APIs limits reproducibility and broad participation. We study whether local open-source ensemble debates can improve alignmentoriented reasoning. Across 150 debates spanning 15 scenarios and five ensemble configurations, ensembles outperform single-model baselines on a 7-point rubric (overall: 3.48 vs. 3.13), with the largest gains in reasoning depth (+19.4%) and argument quality (+34.1%). Improvements are strongest for truthfulness (+1.25 points) and human enhancement (+0.80). We provide code, prompts, and a debate data set, providing an accessible and reproducible foundation for ensemble-based alignment evaluation.

AI对齐模型组可复现性辩论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。