arXiv:2604.27624cs.CLcs.AI2026-04被引 2

构建19万条模拟人类社会属性的对话数据集,研究大模型如何受人格与背景影响表达立场。

Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior

论文配图:Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
图 1 · 摘自论文原文
  • 用19个大模型模拟17种社会属性角色,生成4类敏感议题对话
  • 数据集包含真实情感与立场变化,支持情绪与语义分析
  • 提供交互式平台,可对比不同模型在不同身份下的表达差异

大型语言模型(LLMs)深刻影响社会讨论,但针对其在可控社会与情境提示下输出差异的数据集仍稀缺。认知数字影子(Cognitive Digital Shadows, CDS)是一个包含19万条记录的合成语料库,支持对大模型生成话语的分析。每个记录由19个大模型之一生成,被提示扮演人类人格或AI助手角色。语料库涵盖疫苗/医疗、社交媒体虚假信息、科学领域的性别差距以及STEM刻板印象等4个争议性社会议题。人格化记录编码了17种社会人口学与心理属性,实现大模型提示、语言、立场与推理之间的关联分析。文本经验证具备话题锚定性,可通过可解释NLP(如文本形式思维网络)支持情感分析。CDS配备用户友好的聚合平台与可视化仪表盘,支持按人物、主题和模型进行群体层面的情感与语义框架对比。该提示框架为未来大模型偏见、社会敏感性与对齐性的审计提供了支持。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can strongly shape social discourse, yet datasets investigating how LLM outputs vary across controlled social and contextual prompting remain sparse. Cognitive Digital Shadows (CDS) is a 190,000-record synthetic corpus supporting analyses of LLM-generated discourse. Each CDS record is generated by one of 19 LLMs, prompted to shadow either a human persona or an AI-assistant role. CDS contains LLM responses on 4 controversial societal topics: vaccines/healthcare, social media disinformation, the gender gap in science, and STEM stereotypes. Persona-conditioned records encode 17 sociodemographic and psychological attributes, providing data linking LLMs' prompts, language, stances and reasoning. Texts are validated for topic anchoring and can support emotional analyses via interpretable NLP (e.g. textual forma mentis networks). CDS is enriched by a pooling platform with user-friendly dashboards, enabling easy, interactive group-level comparisons of emotional and semantic framing across personas, topics and models. The CDS prompting framework supports future audits of LLMs' bias, social sensitivity and alignment.

大模型评估社会偏见数字影子话语分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。