通过动态提问提升大模型对个体决策的模拟精度
Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

- 设计三阶段对话框架,动态获取人物背景细节
- 基于追问内容推理的预测准确率达45.5%,优于仅用核心问题的39.3%
- 适合需要高精度个性化决策模拟的研究与应用
大语言模型在模拟特定个体决策时仍面临挑战,主要因人物设定多为静态描述,缺乏价值观、经历和情境线索。本文提出一种自适应访谈框架,通过三阶段对话——核心问题、动态追问与人格总结——收集人物相关信息。利用访谈记录评估模型在道德困境中的决策模拟表现。比较三种对话情境:仅核心问题(Core-10)、完整对话与摘要化人格表征。结果表明,自适应访谈更像一个选择性证据锚定机制:约40%的完整对话路径中,模型实际使用了追问获得的信息;这些基于追问证据的预测准确率(45.5%)显著高于仅依赖核心信息的(39.3%)。说明仅提供丰富背景不足以为效,模型必须真正依据用户特定证据进行决策才可提升精度。
原文摘要 · Abstract (English)
Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information is often provided as static descriptions that miss the values, experiences, and contextual cues needed for individual-level decision simulation. We propose an adaptive interview framework that gathers persona-relevant information through a structured three-stage dialogue: core questions, dynamic follow-ups, and a synthesized personality summary. Using the resulting interview transcripts, we evaluate whether LLMs can simulate participants' decisions in moral dilemma scenarios. We compare three conversational contexts -- Core-10 responses, the full interview dialogue, and a summarized persona representation. We find that adaptive interviewing functions less as a uniform accuracy booster and more as a selective grounding mechanism: follow-up-derived evidence is incorporated in around 40% of full-interview traces, and these follow-up-grounded predictions are more accurate than core-only grounded ones (45.5% vs. 39.3%). These findings highlight that richer persona context alone is insufficient: improvements arise only when models actually ground their decisions in user-specific evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。