arXiv:2506.06113cs.CL2025-06

让大模型学会反映人类对主观任务的分歧,提升社会智能。

Bridging the Gap: In-Context Learning for Modeling Human Disagreement

  • 用上下文学习让模型生成多视角判断,而非单一答案。
  • 零样本下可生成多观点,少样本时仍难覆盖全部分歧。
  • 选例方式比排序方式更重要,需关注标注争议度。

大语言模型在自然语言分类任务中表现优异,但通常依赖聚合标签(如多数投票),会掩盖主观标注中的真实分歧。本研究探讨大模型能否在仇恨言论和冒犯性语言检测等主观任务中捕捉多重视角并反映标注者分歧。采用零样本与少样本上下文学习,评估四种开源大模型在三种标签建模策略下的表现:聚合硬标签,以及分散的硬标签与软标签。在少样本提示中,测试基于文本相似性(BM25、PLM)、标注分歧(熵值)、综合排名及示例排序策略(随机与课程式)。结果表明,零样本下多视角生成可行;而少样本设置常无法覆盖全部人类判断。提示设计与示例选择显著影响性能,但示例排序影响较小。研究揭示了建模主观性的挑战,强调构建更具视角感知力的社会智能模型的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown strong performance on NLP classification tasks. However, they typically rely on aggregated labels-often via majority voting-which can obscure the human disagreement inherent in subjective annotations. This study examines whether LLMs can capture multiple perspectives and reflect annotator disagreement in subjective tasks such as hate speech and offensive language detection. We use in-context learning (ICL) in zero-shot and few-shot settings, evaluating four open-source LLMs across three label modeling strategies: aggregated hard labels, and disaggregated hard and soft labels. In few-shot prompting, we assess demonstration selection methods based on textual similarity (BM25, PLM-based), annotation disagreement (entropy), a combined ranking, and example ordering strategies (random vs. curriculum-based). Results show that multi-perspective generation is viable in zero-shot settings, while few-shot setups often fail to capture the full spectrum of human judgments. Prompt design and demonstration selection notably affect performance, though example ordering has limited impact. These findings highlight the challenges of modeling subjectivity with LLMs and the importance of building more perspective-aware, socially intelligent models.

大模型主观判断上下文学习多视角生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。