arXiv:2603.27057cs.CLcs.AI2026-03中稿 · publication in IEE…

通过提示词注入社会归因知识,减少大模型在社交媒体分析中的偏见。

Debiasing Large Language Models toward Social Factors in Online Behavior Analytics through Prompt Knowledge Tuning

  • 用用户目标和消息上下文构建提示词,引导模型做社会归因推理。
  • 在灾情场景下,零样本分类任务准确率提升,偏见显著降低。
  • 适用于多语言、多灾种的社交媒体行为分析,适合安全敏感场景使用。

归因理论解释了个体如何在社交情境中通过个人因素(特质)和情境因素(外部环境)来解释他人行为。大型语言模型(LLMs)在人类生成语料上训练,可能隐式模仿这一社会归因过程。然而,这些模型在推理中如何运用因果归因仍不明确。尽管链式思维(CoT)等推理范式在各类任务中表现良好,但在推理中忽略社会归因可能导致模型在社交情境中产生偏差。本研究探究将用户意图作为推断特质归因的知识、消息上下文作为推断情境归因的知识,对模型性能的影响。为此,我们提出一种可扩展的方法:基于社交媒体消息的情境与目标,向指令提示中注入两类社会归因提示辅助,以缓解此类偏差。该方法在零样本分类任务中提升了模型性能,并降低了社会归因偏差。我们在两个任务——灾难领域社交媒体中的意图识别与主题检测——上验证了该方法的有效性,覆盖多种灾情类型与多语言环境。实验揭示了Llama3、Mistral和Gemma三款开源模型在社会归因上的固有偏见,并证明了所提策略的改进效果。

原文摘要 · Abstract (English)

Attribution theory explains how individuals interpret and attribute others' behavior in a social context by employing personal (dispositional) and impersonal (situational) causality. Large Language Models (LLMs), trained on human-generated corpora, may implicitly mimic this social attribution process in social contexts. However, the extent to which LLMs utilize these causal attributions in their reasoning remains underexplored. Although using reasoning paradigms, such as Chain-of-Thought (CoT), has shown promising results in various tasks, ignoring social attribution in reasoning could lead to biased responses by LLMs in social contexts. In this study, we investigate the impact of incorporating a user's goal as knowledge to infer dispositional causality and message context to infer situational causality on LLM performance. To this end, we introduce a scalable method to mitigate such biases by enriching the instruction prompts for LLMs with two prompt aids using social-attribution knowledge, based on the context and goal of a social media message. This method improves the model performance while reducing the social-attribution bias of the LLM in the reasoning on zero-shot classification tasks for behavior analytics applications. We empirically show the benefits of our method across two tasks-intent detection and theme detection on social media in the disaster domain-when considering the variability of disaster types and multiple languages of social media. Our experiments highlight the biases of three open-source LLMs: Llama3, Mistral, and Gemma, toward social attribution, and show the effectiveness of our mitigation strategies.

大模型社会归因偏见缓解灾情分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。