用一致性最大化生成无监督价值示例,提升AI对多元人类价值观的对齐效果。
Coherence Maximization Improves Pluralistic Alignment
- 通过最大化标签互预测性,无监督生成指向特定人格的价值示例。
- 一致性高的示例在四个任务中表现媲美人工标注,且泛化能力更强。
- 对预训练数据中少数群体,针对性提问反馈比随机标注更有效。
将AI系统与多样化的个人价值观对齐,需要基于具体示例的价值规范,但无需大量人工干预即可生成此类示例仍是开放挑战。本文研究示例有效性来源,采用内部一致性最大化(ICM)——通过最大化标签间互预测性推断标签——来生成无需人工标注的、针对特定人格的价值示例,引导模型贴近目标群体的价值观。在涵盖分类、偏好和开放式生成的四个基准测试中,ICM推断的上下文示例性能媲美黄金标签。关键发现:一致性的重要性超越单个标签准确率;在准确率不变的前提下,一致性更高的示例泛化能力显著优于不一致的示例。对于预训练数据中代表性不足的人群,针对模型最不确定的问题进行针对性人工反馈,其泛化效果优于同等数量的任意问题标注。结果表明,一致性是可扩展价值规范的关键设计原则,可利用预训练语言模型中已编码的多元人类视角。
原文摘要 · Abstract (English)
Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensive human supervision remains an open challenge. We investigate what makes these examples effective, using Internal Coherence Maximization (ICM) -- which infers labels by maximizing their mutual predictability -- to generate persona-specific examples that steer a model toward a target group's values, without human supervision. Across four benchmarks spanning classification, preference, and open-ended generation, ICM-inferred in-context examples match the performance of gold labels. Crucially, coherence matters beyond individual label accuracy: with accuracy held constant, more coherent examples generalize substantially better than incoherent ones. For personas underrepresented in pretraining data, targeted human feedback on the questions where the model is least certain about a persona's values yields better generalization than the same number of labels on arbitrary questions. These results identify coherence as a key design principle for scalable value specification, leveraging the diverse human perspectives already encoded in pretrained language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。