让对话推荐理解场景变化与隐含偏好,更准更自然。
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation

- 用场景转移估计判断用户是否满意当前环境
- 通过贝叶斯反向推断预测用户对物品的隐含偏好
- 适合做带视觉上下文的智能推荐系统研究者
情境化对话推荐(SCR)结合特定环境中的视觉场景与自然语言对话,提供符合上下文的推荐,贴近真实应用场景。相比传统推荐,SCR需深入理解动态且隐含的用户偏好,因周围环境常影响用户潜在兴趣,且两者随对话不断演变,这显著影响推荐时机与相关性。为此,我们提出情境偏好推理框架(SiPeR),包含两个核心机制:(1)场景转移估计,评估当前场景是否满足用户需求,并在必要时引导用户前往更合适的场景;(2)贝叶斯逆向推理,利用多模态大语言模型(MLLMs)的似然概率,推断用户对场景中候选物品的偏好。在两个代表性基准上的大量实验表明,SiPeR在推荐准确率和响应生成质量上均优于现有方法。代码与数据已公开于 https://github.com/DongdingLin/SiPeR。
原文摘要 · Abstract (English)
Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recommendations, has emerged as a promising research direction due to its close alignment with real-world scenarios. Compared to traditional recommendations, SCR requires a deeper understanding of dynamic and implicit user preferences, as the surrounding scene often influences users' underlying interests, while both may evolve across conversations. This complexity significantly impacts the timing and relevance of recommendations. To address this, we propose situated preference reasoning (SiPeR), a novel framework that integrates two core mechanisms: (1) Scene transition estimation, which estimates whether the current scene satisfies user needs, and guides the user toward a more suitable scene when necessary; and (2) Bayesian inverse inference, which leverages the likelihood of multimodal large language models (MLLMs) to predict user preferences about candidate items within the scene. Extensive experiments on two representative benchmarks demonstrate SiPeR's superiority in both recommendation accuracy and response generation quality. The code and data are available at https://github.com/DongdingLin/SiPeR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。