用少量标注数据让模型更懂多元观点,减少偏见。
Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
- 通过多视角解码和模型引导提升多样性对齐
- 仅用50个样本就优于零样本和少样本基线
- 适合关注公平性与价值观对齐的研究者
随着语言模型对社会影响加深,确保其能反映人类价值观的多样性变得愈发重要。然而,当前主流训练范式通常假设每个问题都有唯一最优答案,导致回应趋于泛化且对齐不足。本文在低资源条件下提出两种方法:多视角解码与模型引导,以增强语言模型的多元对齐能力。实验表明,模型引导在仅使用50个标注样本的情况下,持续优于零样本和少样本基线。所提方法显著降低了仇恨言论检测与虚假信息检测中的误报率,并提升了在GlobalOpinionQA上的分布对齐效果。本工作强调了多样性的重要性,展示了如何让语言模型更好地考虑复杂的人类观点。
原文摘要 · Abstract (English)
As language models have a greater impact on society, it is important to ensure they are aligned to a diverse range of perspectives and are able to reflect nuance in human values. However, the most popular training paradigms for modern language models often assume there is one optimal answer for every query, leading to generic responses and poor alignment. In this work, we aim to enhance pluralistic alignment of language models in a low-resource setting with two methods: pluralistic decoding and model steering. We empirically demonstrate that model steering offers consistent improvement over zero-shot and few-shot baselines with only 50 annotated samples. Our proposed methods decrease false positives in several high-stakes tasks such as hate speech detection and misinformation detection, and improves the distributional alignment to human values in GlobalOpinionQA. We hope our work highlights the importance of diversity and how language models can be adapted to consider nuanced perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。