arXiv:2505.21362cs.CLcs.AI2025-05被引 4

测试大模型如何根据用户年龄教育背景调整回应,发现推理强的模型更稳定。

Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History

  • 用用户资料或对话历史两种方式测试模型适应能力
  • 多数模型随年龄教育变化调整价值表达,但稳定性不一
  • 推理能力强的模型适应更一致,适合需要个性化交互的应用

大语言模型有效互动需适应用户的性别、年龄、职业和教育水平等社会人口学特征。尽管实际应用中常利用对话历史进行上下文理解,但现有评估多集中于单轮提示。本文提出一种评估框架,考察当属性通过用户资料显式引入或通过多轮对话历史隐式体现时,模型行为的一致性。我们构建了一个多智能体合成数据集,将对话历史与不同用户资料配对,并使用价值调查模块(VSM 2013)的问题探测价值表达。结果表明,大多数模型会根据年龄和教育水平的变化调整其表达的价值观,但一致性存在差异;推理能力更强的模型表现出更高的行为一致性,凸显推理在稳健社会人口学适应中的重要性。

原文摘要 · Abstract (English)

Effective engagement by large language models (LLMs) requires adapting responses to users' sociodemographic characteristics, such as age, occupation, and education level. While many real-world applications leverage dialogue history for contextualization, existing evaluations of LLMs' behavioral adaptation often focus on single-turn prompts. In this paper, we propose a framework to evaluate LLM adaptation when attributes are introduced either (1) explicitly via user profiles in the prompt or (2) implicitly through multi-turn dialogue history. We assess the consistency of model behavior across these modalities. Using a multi-agent pipeline, we construct a synthetic dataset pairing dialogue histories with distinct user profiles and employ questions from the Value Survey Module (VSM 2013) (Hofstede and Hofstede, 2016) to probe value expression. Our findings indicate that most models adjust their expressed values in response to demographic changes, particularly in age and education level, but consistency varies. Models with stronger reasoning capabilities demonstrate greater alignment, indicating the importance of reasoning in robust sociodemographic adaptation.

大模型评估社会人口学个性化交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。