通过隐含立场分析,评估模型对多元观点的包容性表现。
EMBRACE: Shaping Inclusive Opinion Representation by Aligning Implicit Conversations with Social Norms
- 用响应立场代理隐含意见,捕捉对话中被忽略的深层观点。
- 在正样本-未标注学习与指令微调模型中验证了规范对齐效果。
- 为构建更公平的对话模型提供可量化的评估框架,适合伦理与社会计算研究者。
构建包容性表达以体现多样性并保障公平参与和价值反映,是众多基于对话的模型的核心目标。然而,现有方法多依赖表面性包含,如提及用户人口统计或社会群体的行为属性,忽视了对话中嵌入的微妙、隐含的意见表达。过度依赖显性线索可能加剧偏差,强化有害或刻板印象。为此,我们重新审视问题,认为公平包容需关注隐含意见表达,并以回应立场验证规范对齐。本研究提出一种对齐评估框架,突出被忽视的隐含对话,评估其与社会规范的一致性。通过将响应立场建模为潜在意见的代理,实现对多元社会观点的审慎与反思性表征。我们采用(i)基于基础分类器的正样本-未标注(PU)在线学习,以及(ii)指令微调的语言模型,评估后训练对齐效果。该框架提供了系统化视角,揭示隐含意见的(误)表征机制,并为实现更包容的模型行为提供路径。
原文摘要 · Abstract (English)
Shaping inclusive representations that embrace diversity and ensure fair participation and reflections of values is at the core of many conversation-based models. However, many existing methods rely on surface inclusion using mention of user demographics or behavioral attributes of social groups. Such methods overlook the nuanced, implicit expression of opinion embedded in conversations. Furthermore, the over-reliance on overt cues can exacerbate misalignment and reinforce harmful or stereotypical representations in model outputs. Thus, we took a step back and recognized that equitable inclusion needs to account for the implicit expression of opinion and use the stance of responses to validate the normative alignment. This study aims to evaluate how opinions are represented in NLP or computational models by introducing an alignment evaluation framework that foregrounds implicit, often overlooked conversations and evaluates the normative social views and discourse. Our approach models the stance of responses as a proxy for the underlying opinion, enabling a considerate and reflective representation of diverse social viewpoints. We evaluate the framework using both (i) positive-unlabeled (PU) online learning with base classifiers, and (ii) instruction-tuned language models to assess post-training alignment. Through this, we provide a principled and structured lens on how implicit opinions are (mis)represented and offer a pathway toward more inclusive model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。