让零售智能体提前预测顾客意图并主动干预,提升服务主动性。
See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

- 用AIDA与BDI建模顾客心理状态,预测行为意图变化。
- 在真实场景中达到0.579的行动分类准确率,优于基线模型。
- 适合研究主动式人机交互与智能导购系统的设计者。
多模态零售智能体不仅需识别顾客行为,还应在明确请求前预判并决定是否干预。本文提出See--Infer--Intervene(SII)框架,要求设备通过观察互动前行为,推断隐含顾客意图,并选择合适服务响应或等待。我们构建了主动意图世界模型(PIWM),以AIDA(注意力、兴趣、欲望、行动)购买阶段和BDI(信念、欲望、意图)心理场表示顾客状态,预测动作条件下的意图转移,并从五类响应中选择:问候、探询、告知、推荐、等待。我们进一步创建GuidanceSalesBench智能零售基准,包含状态记录、互动前视频、候选回应、动作结果及最优动作标签。在使用真实状态条件下,PIWM在30个保留视频上取得0.641的宏F1,优于零样本Qwen2.5-VL-7B基线及缺乏平衡动作监督的训练变体;而仅依赖视频的端到端选择降至0.295,低于5类随机基线0.414,表明视频到状态对齐是部署瓶颈。初步门店试点(20段标注视频)实现0.579的宏F1,另10段视频附带索引级标签已公开。
原文摘要 · Abstract (English)
Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is made. We study this setting through the See--Infer--Intervene (SII) framework, where a device must see pre-interaction behavior, infer latent customer intent, and act by selecting an appropriate service intervention or choosing to wait. We instantiate SII with the Proactive Intent World Model (PIWM), which represents customer state with AIDA (Attention, Interest, Desire, Action) purchasing phases and BDI (belief, desire, intention) psychological fields, predicts action-conditioned intent transitions, and selects from five response classes: Greet, Elicit, Inform, Recommend, and Hold. We further construct GuidanceSalesBench, a smart-retail benchmark containing state manifests, pre-interaction videos, candidate responses, action-conditioned outcomes, and best-action labels. When conditioned on ground-truth customer state to isolate action selection, PIWM achieves 0.641 macro F1 on 30 held-out target videos, outperforming a zero-shot Qwen2.5-VL-7B baseline and training variants without balanced action supervision; end-to-end video-only selection drops to 0.295, below the 5-class balanced random baseline of 0.414, identifying video-to-state grounding as the dominant deployment-time bottleneck. A preliminary staged real-store pilot (recorded with paid participants performing scripted customer behaviors) reaches 0.579 action macro F1 on 20 fully annotated videos, with 10 additional accessible videos released with index-level labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。