预测用户是否需要看图,智能推荐换设备查看商品。
Image-Seeking Intent Prediction for Cross-Device Product Search
- 基于语音查询和商品信息,预测是否需切换到带屏设备看图。
- 在90万条真实语音查询上训练,准确率显著提升。
- 适合做跨设备电商助手的团队参考,提升用户体验。
大型语言模型正重塑电商领域的个性化搜索与用户交互。消费者越来越多地在多种设备间切换,从仅支持语音的助手到多模态显示设备,输入输出能力各异。主动建议切换设备可极大改善体验,但必须精准以避免干扰。本文提出图像需求意图预测(Image-Seeking Intent Prediction),一种面向大模型驱动电商助手的新任务,旨在预判语音商品查询是否应主动触发屏幕设备上的视觉呈现。基于某多设备零售助手的大规模生产数据(含90万条语音查询、关联商品检索结果及图像轮播互动行为信号),我们训练了IRP(Image Request Predictor)模型,利用用户查询语义和检索商品元数据来预测视觉需求。实验表明,结合查询语义与商品数据(尤其是轻量摘要增强后)能持续提升预测准确率;引入可微分的精确度导向损失进一步降低误报。结果表明,大模型可赋能智能、自适应的跨设备购物助手,实现更流畅、个性化的电商体验。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are transforming personalized search, recommendations, and customer interaction in e-commerce. Customers increasingly shop across multiple devices, from voice-only assistants to multimodal displays, each offering different input and output capabilities. A proactive suggestion to switch devices can greatly improve the user experience, but it must be offered with high precision to avoid unnecessary friction. We address the challenge of predicting when a query requires visual augmentation and a cross-device switch to improve product discovery. We introduce Image-Seeking Intent Prediction, a novel task for LLM-driven e-commerce assistants that anticipates when a spoken product query should proactively trigger a visual on a screen-enabled device. Using large-scale production data from a multi-device retail assistant, including 900K voice queries, associated product retrievals, and behavioral signals such as image carousel engagement, we train IRP (Image Request Predictor), a model that leverages user input query and corresponding retrieved product metadata to anticipate visual intent. Our experiments show that combining query semantics with product data, particularly when improved through lightweight summarization, consistently improves prediction accuracy. Incorporating a differentiable precision-oriented loss further reduces false positives. These results highlight the potential of LLMs to power intelligent, cross-device shopping assistants that anticipate and adapt to user needs, enabling more seamless and personalized e-commerce experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。