arXiv:2601.17172cs.CLcs.AI2026-01ACL被引 4

分析大模型生成广告文本时对不同人群的偏见,发现男女老少被区别对待。

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text

  • 用三款主流模型测试生成针对不同性别年龄群体的消息
  • 男性和年轻人的文案更激进进步,女性和老人的更温和传统
  • 真实场景下偏见被放大,适合关注公平性的人参考

大型语言模型(LLMs)正日益具备规模化生成个性化、有说服力文本的能力,引发自动化沟通中的偏见与公平性问题。本文首次系统分析了大模型在执行基于人口统计特征的定向信息生成任务时的表现。我们引入一个受控评估框架,使用GPT-4o、Llama-3.3和Mistral-Large-2.1三款领先模型,在两种生成场景下进行测试:独立生成(隔离内在人口统计影响)与上下文丰富生成(融入主题和区域背景以模拟真实定向)。从词汇内容、语言风格、说服力框架三个维度评估生成消息。以气候传播为例,发现各模型均存在显著的年龄与性别不对称:针对男性和年轻人的文案更强调坚决和进步立场,而针对女性和老年人的则更突出温暖、关怀与传统主题。上下文提示系统性加剧了这些差异,男性目标文案的说服力评分更高,而年龄差异在不同模型间表现不一。结果表明,人口统计刻板印象可能在大模型生成的定向沟通中浮现并强化,凸显了在社会敏感应用中构建具备偏见意识的生成流程与透明审计框架的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly capable of generating personalized, persuasive text at scale, raising new questions about bias and fairness in automated communication. This paper presents the first systematic analysis of how LLMs behave when tasked with demographic-conditioned targeted messaging. We introduce a controlled evaluation framework using three leading models: GPT-4o, Llama-3.3, and Mistral-Large-2.1, across two generation settings: Standalone Generation, which isolates intrinsic demographic effects, and Context-Rich Generation, which incorporates thematic and regional context to emulate realistic targeting. We evaluate generated messages along three dimensions: lexical content, language style, and persuasive framing. We instantiate this framework on climate communication and find consistent age- and gender-based asymmetries across models: male- and youth-targeted messages tend to emphasize more assertive and progressive framing, while female- and senior-targeted messages more often reflect warmth, care, and traditional themes. Contextual prompts systematically amplify these disparities, with persuasion scores being higher for male-targeted messages, while age-related differences vary across models. Our findings demonstrate how demographic stereotypes can surface and intensify in LLM-generated targeted communication, underscoring the need for bias-aware generation pipelines and transparent auditing frameworks that explicitly account for demographic conditioning in socially sensitive applications.

大模型偏见文本生成公平性审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。