arXiv:2604.24536cs.CL2026-04

让AI生成更易被接受的立场妥协,关键在模拟共情中立。

Generating Place-Based Compromises Between Two Points of View

论文配图:Generating Place-Based Compromises Between Two Points of View
图 1 · 摘自论文原文
  • 用观点间共情相似度做迭代反馈,提升妥协质量
  • 实验显示新方法比标准思维链更易被50人接受
  • 训练小模型时直接对齐人类偏好,推理无需共情估计

大型语言模型在学术任务上表现优异,但在社会智能任务(如生成合理妥协)上仍有不足。本文提出一种生成观点间情感中立妥协的方法,通过对比四种提示工程策略,在包含2400个关于公共空间对立观点的数据集上使用Claude 3 Opus进行测试。50名参与者评估了部分生成结果,发现利用妥协与各观点间外部共情相似度作为迭代反馈的策略,显著优于标准链式思维(CoT)推理。结果表明,采用共情中立能有效提升妥协的可接受性。随后,基于生成的妥协数据集,通过基于边距的人类偏好对齐训练了两个小型基础模型,提升了效率并避免了推理阶段的共情估计需求。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel academically but struggle with social intelligence tasks, such as creating good compromises. In this paper, we present methods for generating empathically neutral compromises between two opposing viewpoints. We first compared four different prompt engineering methods using Claude 3 Opus and a dataset of 2,400 contrasting views on shared places. A subset of the gen erated compromises was evaluated for acceptability in a 50-participant study. We found that the best method for generating compromises between two views used external empathic similarity between a compromise and each viewpoint as iterative feedback, outperforming stan dard Chain of Thought (CoT) reasoning. The results indicate that the use of empathic neutrality improves the acceptability of compromises. The dataset of generated compromises was then used to train two smaller foundation models via margin-based alignment of human preferences, improving efficiency and removing the need for empathy estimation during inference.

AI共情立场妥协提示工程偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。