arXiv:2506.14625cs.CLcs.AI2025-06ACL被引 1

多模型协同判断道德难题,提升AI决策一致性。

Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models

  • 用连续评分融合多个模型意见,按可信度加权
  • 对偏差模型进行针对性嵌入优化,降低与共识差异
  • 适用于需要高一致性的伦理决策场景

大型语言模型在道德推理方面表现出色,但在面对复杂多因素的道德困境时,常出现分歧。为此,我们提出一个框架,将多个模型的道德判断整合为集体共识判断,并使明显偏离共识的模型重新对齐。该聚合机制通过连续道德可接受性评分(非二值标签)构建集体概率,依据模型可靠性加权贡献。针对偏离模型,采用目标嵌入优化方法,微调道德哲学理论相关的词元嵌入,最小化与共识的JS散度,同时保持语义完整性。在大规模社会道德困境数据集上的实验表明,该方法能建立稳健的共识并提升单个模型的忠实度。研究凸显了跨模型数据驱动对齐的价值,及其在打造更安全、一致的AI系统中的潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive moral reasoning abilities. Yet they often diverge when confronted with complex, multi-factor moral dilemmas. To address these discrepancies, we propose a framework that synthesizes multiple LLMs' moral judgments into a collectively formulated moral judgment, realigning models that deviate significantly from this consensus. Our aggregation mechanism fuses continuous moral acceptability scores (beyond binary labels) into a collective probability, weighting contributions by model reliability. For misaligned models, a targeted embedding-optimization procedure fine-tunes token embeddings for moral philosophical theories, minimizing JS divergence to the consensus while preserving semantic integrity. Experiments on a large-scale social moral dilemma dataset show our approach builds robust consensus and improves individual model fidelity. These findings highlight the value of data-driven moral alignment across multiple models and its potential for safer, more consistent AI systems.

道德推理多模型融合嵌入优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。