让大模型学会预判用户状态变化,提升对话长期效果。
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation

- 基于自由能原理构建用户状态模型,动态追踪隐藏心理状态。
- 在医疗咨询任务中,长对话效果显著提升,状态预测更准确。
- 适合需要深度理解用户演变的个性化对话系统开发者。
个性化对话不仅需记忆用户显式历史,还需推断随交互演化的隐含用户状态,以制定恰当响应策略。现有基于记忆或资料的方法主要依赖可观测信息,难以建模用户状态动态或根据行动对未来状态的影响做选择。本文提出PUMA(面向行动选择的前瞻性用户状态建模)框架,基于自由能原理(FEP),将个性化视为部分可观测下的决策问题,核心是显式建模隐含用户状态及其受动作影响的演化规律。每轮对话中,PUMA维护用户隐状态的信念分布,更新观测生成与动作条件下的状态转移模型,并通过最小化预期自由能选择对话行为,在认知目标与实用目标间取得平衡。该方法将个性化从被动记忆检索转向对用户演化过程的模型驱动决策。我们在面向医疗咨询与动机访谈的基准上实例化PUMA,采用带隐状态标注的数据集进行严格评估。实验表明,PUMA在长时对话中表现更优,同时保持高质量回复;跨数据集研究显示其用户状态估计与未来状态预测更具鲁棒性。
原文摘要 · Abstract (English)
Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interaction and shape appropriate response strategies. Existing memory- and profile-based methods primarily reuse observable user information, offering limited support for modeling user-state dynamics or selecting actions based on how they shape future user states. We propose PUMA (Prospective User-state Modeling for Action selection), a framework grounded in the Free Energy Principle (FEP) that formulates personalization as decision-making under partial observability, centered on an explicit user state model that captures latent user states and their action-conditioned dynamics. At each turn, PUMA maintains a belief over the user's hidden state, refines the user state model for observation generation and action-conditioned state transition, and selects dialogue actions by minimizing expected free energy, balancing epistemic and pragmatic objectives under a unified criterion. This formulation shifts personalization from passive memory retrieval to model-based decision-making over user evolution. We instantiate PUMA on healthcare-oriented counseling and motivational interviewing benchmarks with latent state annotations for rigorous evaluation. Experiments show that PUMA improves long-horizon dialogue outcomes while maintaining strong response quality, and a cross-dataset study demonstrates more reliable user-state estimation and next-state prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。