arXiv:2411.10912cs.CL2024-11ACL被引 8

让AI在推理时根据不同群体偏好精准选择示例,提升对齐效果。

SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment

  • 构建情景库与群体感知检索机制,识别跨群体价值差异。
  • 实验显示最高提升0.16分(5分制),且各群体收益更均衡。
  • 适合需要兼顾多元价值观的AI对齐场景,如公共政策生成。

当不同群体的价值观存在差异时,一种模型对齐方法是在推理阶段引导模型贴近特定群体的偏好。然而,现有基于上下文学习的技术仅考虑示例相似性,未关注群体间的价值差异。本文提出SPICA框架,在上下文示例检索中引入群体级差异考量。SPICA包含三大设计:情景库、群体感知检索度量和上下文对齐提示。在涵盖四个社会人口群体(n=544)的对齐任务评估中,该方法能更准确地检索出符合实际偏好的示例;最优提示配置采用多对比响应展示示例。在端到端评估中(n=120),SPICA获得更高评分,群体最高提升达+0.16分(5分制)。此外,其增益分布更均匀,所有群体均受益,而非仅部分群体。最后发现,尽管忽略群体差异的方法可对齐聚合值,但不适合价值分歧明显的群体。

原文摘要 · Abstract (English)

When different groups' values differ, one approach to model alignment is to steer models at inference time towards each group's preferences. However, techniques like in-context learning only consider similarity when drawing few-shot examples and not cross-group differences in values. We propose SPICA, a framework that accounts for group-level differences during in-context example retrieval. SPICA introduces three designs: scenario banks, group-informed retrieval metrics, and in-context alignment prompts. From an evaluation of SPICA on an alignment task collecting inputs from four demographic groups ($n = 544$), our metrics retrieve in-context examples that more closely match observed preferences, with the best prompt configuration using multiple contrastive responses to demonstrate examples. In an end-to-end evaluation ($n = 120$), we observe that SPICA is higher rated than similarity-based retrieval, with groups seeing up to a +0.16 point improvement on a 5 point scale. Additionally, gains from SPICA were more uniform, with all groups benefiting from alignment rather than only some. Finally, we find that while a group-agnostic approach can align to aggregated values, it is not most suited for divergent groups.

模型对齐群体差异上下文学习公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。