多模型协同判断道德难题,提升AI决策一致性。
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
- 用连续评分融合多个模型意见,按可信度加权
- 对偏差模型进行针对性嵌入优化,降低与共识差异
- 适用于需要高一致性的伦理决策场景
大型语言模型在道德推理方面表现出色,但在面对复杂多因素的道德困境时,常出现分歧。为此,我们提出一个框架,将多个模型的道德判断整合为集体共识判断,并使明显偏离共识的模型重新对齐。该聚合机制通过连续道德可接受性评分(非二值标签)构建集体概率,依据模型可靠性加权贡献。针对偏离模型,采用目标嵌入优化方法,微调道德哲学理论相关的词元嵌入,最小化与共识的JS散度,同时保持语义完整性。在大规模社会道德困境数据集上的实验表明,该方法能建立稳健的共识并提升单个模型的忠实度。研究凸显了跨模型数据驱动对齐的价值,及其在打造更安全、一致的AI系统中的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive moral reasoning abilities. Yet they often diverge when confronted with complex, multi-factor moral dilemmas. To address these discrepancies, we propose a framework that synthesizes multiple LLMs' moral judgments into a collectively formulated moral judgment, realigning models that deviate significantly from this consensus. Our aggregation mechanism fuses continuous moral acceptability scores (beyond binary labels) into a collective probability, weighting contributions by model reliability. For misaligned models, a targeted embedding-optimization procedure fine-tunes token embeddings for moral philosophical theories, minimizing JS divergence to the consensus while preserving semantic integrity. Experiments on a large-scale social moral dilemma dataset show our approach builds robust consensus and improves individual model fidelity. These findings highlight the value of data-driven moral alignment across multiple models and its potential for safer, more consistent AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。