arXiv:2510.00174cs.CL2025-10EMNLP被引 3

构建包含观点与信念解释的用户数据集,助力语言模型个性化

PrimeX: A Dataset of Worldview, Opinion, and Explanation

  • 收集858名美国居民的观点、信念解释和世界观数据
  • 信念解释和世界观信息显著提升观点预测效果
  • 适用于自然语言处理与心理学交叉研究

随着语言模型应用深入,对用户个体特征的刻画需求日益增强。本研究聚焦观点预测任务,构建了名为PrimeX的数据集,包含858名美国居民的公开意见调查数据,以及受访者对特定观点的书面解释和普里马尔世界观问卷(Primal World Belief survey)结果。我们对数据进行了初步分析,验证了信念解释和世界观信息在提升模型个性化能力方面的价值。结果表明,这些额外的信念信息能有效增强语言模型对个体态度的理解,为自然语言处理与心理学研究提供新范式。

原文摘要 · Abstract (English)

As the adoption of language models advances, so does the need to better represent individual users to the model. Are there aspects of an individual's belief system that a language model can utilize for improved alignment? Following prior research, we investigate this question in the domain of opinion prediction by developing PrimeX, a dataset of public opinion survey data from 858 US residents with two additional sources of belief information: written explanations from the respondents for why they hold specific opinions, and the Primal World Belief survey for assessing respondent worldview. We provide an extensive initial analysis of our data and show the value of belief explanations and worldview for personalizing language models. Our results demonstrate how the additional belief information in PrimeX can benefit both the NLP and psychological research communities, opening up avenues for further study.

用户建模观点预测信念数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。