用人格化提示让AI模仿人类的快慢思维,更像真人推理。
Giving AI Personalities Leads to More Human-Like Reasoning
- 基于五大性格模型设计人格化提示,引导AI生成不同风格的回答。
- 优化后的人格提示使开源模型(如Llama)预测人类回答分布更准确。
- 适合研究人类认知、AI可解释性或具身智能的学者与开发者。
在计算认知建模中,捕捉人类判断与决策的完整谱系(而不仅是最优行为)是一项重大挑战。本研究探索大语言模型(LLMs)是否可通过预测直观快速的系统1和深思熟虑的系统2过程,模拟人类推理的多样性。我们设计了基于自然语言蕴含(NLI)新变体的推理任务,以诱发系统1与系统2反应。通过众包收集人类响应,并建模整个分布而非仅多数答案。采用受大五人格模型启发的性格化提示,引导AI输出体现特定人格特质的反应,从而捕捉人类推理多样性。结合遗传算法优化提示权重,该方法在传统机器学习模型对比下表现优异。结果显示,开放源代码模型(如Llama、Mistral)在模拟人类响应分布上优于专有GPT模型。性格化提示配合遗传算法显著提升模型对人类反应分布的预测能力,表明捕捉非最优、自然化的推理需融合多样推理风格与心理画像。研究结论认为,人格化提示结合遗传算法是增强AI‘类人’推理的有力路径。
原文摘要 · Abstract (English)
In computational cognitive modeling, capturing the full spectrum of human judgment and decision-making processes, beyond just optimal behaviors, is a significant challenge. This study explores whether Large Language Models (LLMs) can emulate the breadth of human reasoning by predicting both intuitive, fast System 1 and deliberate, slow System 2 processes. We investigate the potential of AI to mimic diverse reasoning behaviors across a human population, addressing what we call the "full reasoning spectrum problem". We designed reasoning tasks using a novel generalization of the Natural Language Inference (NLI) format to evaluate LLMs' ability to replicate human reasoning. The questions were crafted to elicit both System 1 and System 2 responses. Human responses were collected through crowd-sourcing and the entire distribution was modeled, rather than just the majority of the answers. We used personality-based prompting inspired by the Big Five personality model to elicit AI responses reflecting specific personality traits, capturing the diversity of human reasoning, and exploring how personality traits influence LLM outputs. Combined with genetic algorithms to optimize the weighting of these prompts, this method was tested alongside traditional machine learning models. The results show that LLMs can mimic human response distributions, with open-source models like Llama and Mistral outperforming proprietary GPT models. Personality-based prompting, especially when optimized with genetic algorithms, significantly enhanced LLMs' ability to predict human response distributions, suggesting that capturing suboptimal, naturalistic reasoning may require modeling techniques incorporating diverse reasoning styles and psychological profiles. The study concludes that personality-based prompting combined with genetic algorithms is promising for enhancing AI's 'human-ness' in reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。