arXiv:2508.15830cs.CLcs.AI2025-08被引 11

检测大模型能否从中性问题推断用户人口属性,揭示潜在隐私风险。

DAIQ: Auditing Demographic Attribute Inference from Question in LLMs

  • 构建新框架DAIQ,评估模型在不确定情境下是否过度推断人口属性。
  • 18个模型在6个领域中均会从普通问题推断出性别、年龄等属性。
  • 提示词调整可显著减少误推断,无需微调模型。

当前对大语言模型的社会偏见评估多依赖明确提及人口属性的提示,忽视了模型是否能从中性问题中推断敏感属性。这种推断属于认知越界,引发隐私担忧。本文提出DAIQ框架,用于诊断在认知不确定性下的人口属性推断行为。我们在六个真实场景、五个人口属性上评估了18个开源与闭源模型,发现多个模型会从中性问题中推断出性别、年龄等属性,且倾向默认社会主流类别,生成符合刻板印象的推理过程。此类行为在不同模型家族、规模及解码设置下持续存在,表明其依赖于预训练数据中的群体先验。我们还发现推断结果会影响下游输出,而采用规避导向提示可显著降低非预期推断,无需模型微调。结果表明现有偏见评估不完整,需建立新标准,不仅考察模型如何回应人口信息,更要评估其是否应推断这些信息。

原文摘要 · Abstract (English)

Recent evaluations of Large language models (LLMs) audit social bias primarily through prompts that explicitly reference demographic attributes, overlooking whether models infer sensitive demographics from neutral questions. Such inference constitutes epistemic overreach and raises concerns for privacy. We introduce Demographic Attribute Inference from Questions (DAIQ), a diagnostic audit framework for evaluating demographic inference under epistemic uncertainty. We evaluate 18 open- and closed-source LLMs across six real-world domains and five demographic attributes. We find that many models infer demographics from neutral questions, defaulting to socially dominant categories and producing stereotype-aligned rationales. These behaviors persist across model families, scales and decoding settings, indicating reliance on learned population priors. We further show that inferred demographics can condition downstream responses and that abstention oriented prompting substantially reduces unintended inference without model fine-tuning. Our results suggest that current bias evaluations are incomplete and motivate evaluation standards that assess not only how models respond to demographic information, but whether they should infer it at all.

大模型安全隐私风险偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。