arXiv:2606.07123cs.CL2026-06

让AI理解语言背后的社会视角差异,基于2.8万条标注数据建模不同群体的解读变化。

Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings

论文配图:Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings
图 1 · 摘自论文原文
  • 融合文本与人口统计特征构建联合嵌入,捕捉社会视角多样性
  • 相比纯文本模型,宏平均PR-AUC提升5.9%-6.5%且显著优于基线
  • 验证了人口统计信息具有真实预测能力,非虚假关联

语言中的社会意义本质上具有视角性,随标注者背景、人口统计特征和意识形态而异。然而,大多数NLP系统将这种差异压缩为单一真值标签,忽略了多元解释。本文在包含2.8万条人工标注的数据集上,沿视角光谱建模社会维度,评估零样本、少样本与微调等多种范式,提出融合嵌入方法,整合文本与人口统计表征。所有融合策略均显著优于仅文本基线(相对宏平均PR-AUC提升5.9%-6.5%),随机打乱实验确认人口统计特征携带真实预测信号而非虚假相关。

原文摘要 · Abstract (English)

Social meaning in language is inherently perspectival, varying across annotator backgrounds, demographics, and ideological positions. However, most NLP systems collapse this variation into a single ground-truth label, ignoring the diversity of interpretations. In this work, we model social dimensions along a perspectivist spectrum, capturing how interpretations vary across demographic groups on a dataset consisting of 28k human annotations. We benchmark multiple modeling paradigms, including zero-shot, few-shot, and fine-tuned approaches, and propose fusion embeddings that integrate textual and demographic representations. Our fusion models yield consistent and statistically significant improvements over text-only baselines across all fusion strategies (+5.9-6.5% relative macro PR-AUC), with shuffle ablations confirming that demographic profiles carry genuine predictive signal rather than spurious correlations.

社会语义视角建模融合嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。