让AI在推理时根据不同群体偏好精准选择示例,提升对齐效果。
SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment
- 构建情景库与群体感知检索机制,识别跨群体价值差异。
- 实验显示最高提升0.16分(5分制),且各群体收益更均衡。
- 适合需要兼顾多元价值观的AI对齐场景,如公共政策生成。
当不同群体的价值观存在差异时,一种模型对齐方法是在推理阶段引导模型贴近特定群体的偏好。然而,现有基于上下文学习的技术仅考虑示例相似性,未关注群体间的价值差异。本文提出SPICA框架,在上下文示例检索中引入群体级差异考量。SPICA包含三大设计:情景库、群体感知检索度量和上下文对齐提示。在涵盖四个社会人口群体(n=544)的对齐任务评估中,该方法能更准确地检索出符合实际偏好的示例;最优提示配置采用多对比响应展示示例。在端到端评估中(n=120),SPICA获得更高评分,群体最高提升达+0.16分(5分制)。此外,其增益分布更均匀,所有群体均受益,而非仅部分群体。最后发现,尽管忽略群体差异的方法可对齐聚合值,但不适合价值分歧明显的群体。
原文摘要 · Abstract (English)
When different groups' values differ, one approach to model alignment is to steer models at inference time towards each group's preferences. However, techniques like in-context learning only consider similarity when drawing few-shot examples and not cross-group differences in values. We propose SPICA, a framework that accounts for group-level differences during in-context example retrieval. SPICA introduces three designs: scenario banks, group-informed retrieval metrics, and in-context alignment prompts. From an evaluation of SPICA on an alignment task collecting inputs from four demographic groups ($n = 544$), our metrics retrieve in-context examples that more closely match observed preferences, with the best prompt configuration using multiple contrastive responses to demonstrate examples. In an end-to-end evaluation ($n = 120$), we observe that SPICA is higher rated than similarity-based retrieval, with groups seeing up to a +0.16 point improvement on a 5 point scale. Additionally, gains from SPICA were more uniform, with all groups benefiting from alignment rather than only some. Finally, we find that while a group-agnostic approach can align to aggregated values, it is not most suited for divergent groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。