增加用户属性提示反而降低大模型预测准确率,关键在属性质量而非数量。
Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

- 用不同数量和组合的用户属性作提示,测试大模型表现
- 1-3个高质量属性时模型与人工标注一致性最高,全属性反而下降
- 属性信号的可学习性和方向一致性决定提示效果,适合做个性化模型调优
我们研究了在五个任务中,将标注者人口统计学属性作为提示线索时,如何影响大语言模型(LLM)预测与人工标注的一致性。采用五种开源大模型,系统地改变提示中人口统计学成分的数量与组合,涵盖从单属性到全属性的所有配置。实验揭示三个核心发现:第一,在包含一至三个高信号属性时,一致性达到峰值,而在完整属性集下则下降,确立了明显的过拟合阈值;第二,人类标注受人口统计学影响的整体强度,并不能预测哪些属性能提升模型对齐度;相反,必须同时考虑每个属性标注信号的可学习性与方向一致性;第三,神经元探针显示,只有在标注信号一致时,特定激活才与对齐提升相关,仅激活量大并不意味着可调控。这些结果表明,人口统计学提示并非单一干预手段,其有效性高度依赖于属性信号质量、任务特征及模型架构。
原文摘要 · Abstract (English)
We investigate how annotator demographic attributes, supplied as prompt cues, shape the alignment between large language model (LLM) predictions and human annotations across five tasks. Using five open-source LLMs, we systematically vary the number and composition of demographic components in the prompt, spanning every combination from single-attribute through full-attribute configurations. Our experiments reveal three principal findings. First, alignment consistently peaks with one to three high-signal attributes and degrades under the full attribute set, establishing a clear over-specification threshold. Second, the overall magnitude of demographic influence on human annotations does not predict which attributes improve LLM alignment; instead, both the learnability and the directional coherence of each attribute's annotation signal need to be considered jointly. Third, neuron probing reveals that specialized activation correlates with alignment gains only under coherent annotation signals, and that activation volume alone does not imply steerability. Together, these results demonstrate that demographic prompting is not a monolithic intervention: its utility is highly context-dependent, shaped by attribute signal quality, task characteristics, and model architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。