通过多层组合融合,让AI模型更懂不同情境下的道德判断。
Contextual Value Alignment via Multilayer Combinatorial Fusion

- 用多个道德代理模拟不同价值观,再组合融合输出
- 在标准评测中优于单代理和单层融合方法
- 适合需要尊重多元伦理的AI应用开发
大语言模型与人类价值观对齐仍是可信AI的核心挑战。现有方法如RLHF、CAI等多依赖单一代理和统一奖励机制,难以体现伦理多样性、情境适应性及多主体道德推理动态。本文提出多层组合融合情境价值对齐框架(MCF-CVA),首先构建多个分别代表不同价值观的道德代理,通过评分与排序组合、平均与加权聚合等方式进行组合扩展,再经缩减回原数量,形成多层迭代的扩展-缩减(EAR)过程,直至满足停止条件。该框架在欧几里得评分空间与凯门尼排序空间双架构下实现,利用代理间认知多样性缓解冲突与冗余,生成更贴合上下文的人类价值观响应。实验表明,该框架在标准指标上超越单代理基线、单层多代理结果及以往聚合方法,证明其在提升大模型情境价值对齐方面具有鲁棒性与有效性。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified reward system. This limits their ability to capture ethical pluralism, adapt to diverse moral contexts, and reflect the dynamics of multi-agent moral reasoning. In this work, we propose a framework that utilizes multilayer combinatorial fusion for contextual value alignment (MCF-CVA). At the first layer of the framework, it instantiates multiple moral agents, each fine-tuned to represent a distinctive value. Their outputs are then expanded combinatorially using both score- and rank-combinations as well as average and weighted aggregations. These combined models are then reduced to the same number of initial moral agents. This expansion and reduction (EAR) process continues for multi-layers until a stopping criterion is reached. The MCF-CVA framework leverages cognitive diversity between agents to mitigate conflicts and redundancies across multiple agents, producing responses that better reflect contextual human values. The framework using the EAR algorithm is performed on the dual architecture of Euclidean score space and Kemeny rank space. Empirical evaluations demonstrated that the proposed framework outperforms single-agent baselines, multi-agent single-layer results, and previous aggregation approaches on standard metrics, showing that the MCF-CVA framework provides a robust and effective mechanism for advancing contextual value alignment in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。