用逆向心理模型理解用户行为背后的动机,提升推荐精准度。
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

- 从用户行为反推其信念与决策特征,构建动态用户画像。
- 在四个任务中表现优于或接近真实人格评估结果。
- 适用于生成式界面、跨模态应用,适合智能交互研究者。
现代推荐系统将用户行为视作偏好可靠指标,但实际交互常反映探索或比较行为而非稳定偏好表达。随着界面从静态布局演进至生成式UI和沉浸式扩展现实(XR),亟需更深层、跨模态的用户理解:自适应环境不仅需决定展示内容,还需判断时机、位置、显著性,以及用户行为背后的动因。本文提出逆向心理理论(Inverse Theory of Mind, IToM)流程,从观察到的交互行为反推解释其背后的心理状态,包括用户所选项及可选项,利用大语言模型进行反事实推理生成基于证据的自然语言信念陈述,并通过多假设归纳推理整合为结构化用户人格画像。在OPeRA数据集上,针对真实人格评估、态度调查和访谈构建的人格画像,四项任务测试显示:推断出的人格画像匹配或超越真实人格,且多假设推理对准确预测人格至关重要。进一步通过基于人格的VisionOS空间银行应用验证了跨模态迁移能力。
原文摘要 · Abstract (English)
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。