用户性格影响对大模型的偏好,不同性格更倾向特定模型。
Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
- 按人格类型分组测试,发现性格决定模型偏好
- 理性型偏爱GPT-4,理想型偏好Claude 3.5
- 揭示传统评估忽略的人格-模型匹配差异
随着大语言模型(LLMs)融入日常协作流程,用户通过多轮交互共同塑造结果,一个关键问题浮现:不同人格特质的用户是否会系统性地偏好某些模型?本研究招募32名参与者,按四种凯尔西人格类型均衡分配,评估其与GPT-4和Claude 3.5在四项协作任务中的表现:数据分析、创意写作、信息检索与写作辅助。结果表明,人格驱动的偏好显著:理性型用户在目标导向任务中强烈偏好GPT-4,而理想型用户在创意与分析任务中更倾向Claude 3.5。其他类型表现出任务依赖性偏好。定性反馈的情感分析验证了这些模式。值得注意的是,总体帮助度评分在模型间相似,凸显人格分析能揭示传统评估所忽视的模型差异。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) increasingly integrate into everyday workflows, where users shape outcomes through multi-turn collaboration, a critical question emerges: do users with different personality traits systematically prefer certain LLMs over others? We conducted a study with 32 participants evenly distributed across four Keirsey personality types, evaluating their interactions with GPT-4 and Claude 3.5 across four collaborative tasks: data analysis, creative writing, information retrieval, and writing assistance. Results revealed significant personality-driven preferences: Rationals strongly preferred GPT-4, particularly for goal-oriented tasks, while idealists favored Claude 3.5, especially for creative and analytical tasks. Other personality types showed task-dependent preferences. Sentiment analysis of qualitative feedback confirmed these patterns. Notably, aggregate helpfulness ratings were similar across models, showing how personality-based analysis reveals LLM differences that traditional evaluations miss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。