arXiv:2507.00439cs.CL2025-07ACL被引 1

用简单监督提升大模型对不同人群的主观认知对齐能力

Improving the Distributional Alignment of LLMs using Supervision

  • 通过添加监督信号优化模型输出分布与多元人群的匹配度
  • 在三大数据集上验证,显著提升跨群体对齐效果
  • 为后续研究提供可复现的评测基准,适合对齐研究者参考

准确对齐大语言模型与多元人群在主观问题上的观点具有重要意义。本文表明,加入简单监督可更一致地改善模型生成分布与不同人群之间的对齐程度,该结论在涵盖公共卫生、公众意见及价值观信念的三个数据集上得到验证。除评估平均对齐水平外,还分析了对齐性能在特定人群间的差异。研究覆盖多种大模型和提示策略,提供了推动未来研究的基准评测体系。

原文摘要 · Abstract (English)

The ability to accurately align LLMs with diverse population groups on subjective questions would have great value. In this work, we show that adding simple supervision can more consistently improve the alignment of LLM-generated distributions with diverse population groups, as measured across three datasets spanning public health, public opinion, and values and beliefs. Beyond evaluating average alignment, we also report how alignment varies across specific groups. Our broad findings provide insights into the distributional alignment of LLM generations with diverse populations. By conducting evaluation over many LLMs and prompting strategies, we provide a benchmark to stimulate future research.

大模型对齐主观判断监督学习群体差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。