构建首个真实时尚对话推荐数据集,支持多模态个性化推荐研究。
VOGUE: A Multimodal Dataset for Conversational Recommendation in Fashion
- 收集60场真人对话,每场含35条精细标注语句,覆盖真实购物场景。
- 包含用户画像、商品元数据和双端评分,可评估推荐准确性与满意度。
- 揭示视觉引导对话的阶段性特征,挑战现有大模型泛化能力。
多模态对话推荐正成为通过自然对话实现个性化体验的新兴范式,其优势在于融合视觉与上下文信息。然而,现有数据集仍存在局限:部分为模拟对话,缺少用户历史记录或缺乏充分反馈,限制了研究深度。为此,我们推出VOGUE数据集,包含60场真人对话,共2100条细粒度标注语句,基于真实时尚购物场景。每段对话配有共享视觉目录、商品元数据、用户时尚档案及对话后用户(Seekers)与推荐者(Assistants)的评分。该设计支持对对话推理能力的严格评估,不仅涵盖预测偏好与真实偏好的对齐,还包括对完整评分分布的校准,以及显式与隐式满意度信号的对比分析。对VOGUE的分析揭示出视觉引导对话的独特动态,如推荐者常以特征组形式同时推荐商品,形成由用户批评与修正连接的对话阶段。基准测试显示,尽管多模态大语言模型在整体偏好对齐上接近人类水平,但在还原人类评分分布方面存在系统性误差,且难以将偏好推断推广至未明确讨论的商品。这些发现使VOGUE成为研究多模态对话系统的重要资源,并构成对当前顶级多模态基础模型(如GPT-5-mini与Gemini-2.5-Flash)的挑战。
原文摘要 · Abstract (English)
Multimodal conversational recommendation has recently emerged as a promising paradigm for delivering personalized experiences through natural dialogue enriched by visual and contextual grounding. Yet currently available multimodal conversational recommendation datasets remain limited: existing resources either simulate conversations, omit user history or fail to collect sufficiently detailed feedback, which constrain the types of research and evaluation they support. To address these gaps we introduce VOGUE, a dataset of 60 human human dialogues containing 2100 granularly labeled utterances in realistic fashion shopping scenarios. Each dialogue is paired with a shared visual catalogue, item metadata, user fashion profiles and post conversation ratings from both users (Seekers) and recommenders (Assistants). This design enables rigorous evaluation of conversational inference, including not only alignment between predicted and ground truth preferences but also calibration against full rating distributions and comparison with explicit and implicit user satisfaction signals. Our analyses of VOGUE reveal distinctive dynamics of visually grounded dialogue, e.g. recommenders frequently recommend items simultaneously in feature based groups, which creates distinct conversational phases bridged by Seeker critiques and refinements. Benchmarking Multimodal Large Language Models against human Recommenders shows that while MLLMs approach human level alignment in aggregate they exhibit systematic distribution errors in reproducing human ratings and struggle to generalize preference inference beyond explicitly discussed items. These findings establish VOGUE as both a unique resource for studying multimodal conversational systems and a challenge dataset beyond the current recommendation capabilities of existing top tier multimodal foundation models such as GPT-5-mini and Gemini-2.5-Flash.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。