arXiv:2603.11126cs.MAcs.CL2026-03中稿 · 2026 IEEE Internat…被引 2

用多个道德代理融合判断,让大模型更符合人类多元价值观。

Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion

  • 构建多个代表不同伦理立场的智能体,通过组合融合分析整合意见。
  • 在标准评测中优于单智能体和传统聚合方法,提升对齐效果。
  • 适合关注AI安全、伦理对齐的研究者与开发者参考。

将大语言模型(LLMs)与人类价值观对齐是确保其可信与安全部署的核心挑战。现有方法如基于人类反馈的强化学习(RLHF)及其变体通常依赖单一评估者或狭义奖励信号,难以捕捉伦理多样性。本文提出价值对齐系统——组合融合分析(VAS-CFA),通过实例化多个经过微调以代表不同规范性视角的道德代理,并采用基于排名与评分的组合融合分析(CFA)聚合其输出。该设计利用智能体间的认知多样性,缓解冲突与冗余,生成更贴近人类价值观的回应。实证评估表明,VAS-CFA在标准指标上超越单智能体基线及以往聚合方法,证明多智能体融合是提升LLM价值对齐的有效机制。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human values is a central challenge for ensuring trustworthy and safe deployment. While existing methods such as Reinforcement Learning from Human Feedback (RLHF) and its variants have improved alignment, they often rely on a single evaluator or narrowly defined reward signals, limiting their ability to capture ethical pluralism. In this work, we propose the Value Alignment System using Combinatorial Fusion Analysis (VAS-CFA), a framework that operationalizes multi-agent fusion alignment. It instantiates multiple moral agents, each fine-tuned to represent a distinct normative perspective, and fuses their outputs using CFA with both rank- and score-based aggregation. This design leverages cognitive diversity, between agents, to mitigate conflicts and redundancies across multiple agents, producing responses that better reflect human values. Empirical evaluation demonstrates that VAS-CFA outperforms both single agent baselines and prior aggregation approaches on standard metrics, showing that multi-agent fusion provides a robust and effective mechanism for advancing value alignment in LLMs.

价值对齐多智能体伦理安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。