通过分阶段警惕与间隔通信,提升多智能体辩论中价值观对齐效果。
Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates
- 智能体分层评估风险,按需通信,减少冗余交互。
- 在多个数据集上显著降低有害内容生成,尤其擅长防欺诈。
- 适用于不同规模模型和任务,兼容对齐与未对齐基模。
近年来,大语言模型在满足多样化人类需求方面表现优异,但其训练数据可能引入有害内容,凸显价值对齐的重要性。主流依赖反馈学习和监督训练的方法资源消耗大,且可能限制模型潜力。多智能体辩论(MAD)通过智能体间交互生成可靠答案,提供更高效创新的解决方案。为将MAD应用于价值对齐,本文分析了辩论结果的有用性与无害性及个体回应的关系,提出基于MAD的渐进式警惕与间隔通信(GVIC)框架。GVIC使智能体能以不同警惕级别评估风险,并通过间隔通信交换多样信息。理论证明GVIC在优化辩论效率的同时降低通信开销。实验表明,GVIC在多种任务和数据集上持续优于基线方法,尤其在有害内容缓解和欺诈防范方面表现突出。此外,GVIC在不同基模型大小、包括对齐与未对齐模型以及各类任务中均表现出强适应性。
原文摘要 · Abstract (English)
In recent years, large language models have shown exceptional performance in fulfilling diverse human needs. However, their training data can introduce harmful content, underscoring the necessity for robust value alignment. Mainstream methods, which depend on feedback learning and supervised training, are resource-intensive and may constrain the full potential of the models. Multi-Agent Debate (MAD) offers a more efficient and innovative solution by enabling the generation of reliable answers through agent interactions. To apply MAD to value alignment, we examine the relationship between the helpfulness and harmlessness of debate outcomes and individual responses, and propose a MAD based framework Gradual Vigilance and Interval Communication (GVIC). GVIC allows agents to assess risks with varying levels of vigilance and to exchange diverse information through interval communication. We theoretically prove that GVIC optimizes debate efficiency while reducing communication overhead. Experimental results demonstrate that GVIC consistently outperforms baseline methods across various tasks and datasets, particularly excelling in harmfulness mitigation and fraud prevention. Additionally, GVIC exhibits strong adaptability across different base model sizes, including both unaligned and aligned models, and across various task types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。