arXiv:2508.18132cs.IRcs.AI2025-08被引 1

让生成式推荐系统在对话中动态调整检索结果,更懂用户不断变化的需求。

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

  • 引入测试时重排序机制,在对话过程中迭代优化检索结果。
  • 多轮评测显示平均提升14.5点MRR、10.6点nDCG@1。
  • 特别适合复杂多轮商品推荐场景,对模糊或演变的用户意图更敏感。

电商快速发展暴露了传统商品检索系统在处理复杂多轮用户交互方面的局限。近年来,基于多模态大语言模型(MLLM)的生成式检索展现出潜力,但多数方法仅适配单轮场景,难以捕捉多轮对话中用户意图的演化与迭代特性。同时,测试时扩展(Test-Time Scaling)虽能通过推理时迭代优化提升大模型性能,但其有效性依赖于明确的问题空间和模型自纠错能力——这些条件在对话式商品搜索中通常不成立。用户查询常模糊且动态变化,而MLLM难以在固定商品库中准确定位。为此,我们提出一种新框架,将测试时扩展引入多模态对话推荐。该框架基于生成式检索器,进一步引入测试时重排序(TTR)机制,持续提升检索精度并更好匹配对话中的演化意图。多个基准测试表明,该方法实现平均14.5点的MRR提升和10.6点的nDCG@1提升。

原文摘要 · Abstract (English)

The rapid evolution of e-commerce has exposed the limitations of traditional product retrieval systems in managing complex, multi-turn user interactions. Recent advances in multimodal generative retrieval -- particularly those leveraging multimodal large language models (MLLMs) as retrievers -- have shown promise. However, most existing methods are tailored to single-turn scenarios and struggle to model the evolving intent and iterative nature of multi-turn dialogues when applied naively. Concurrently, test-time scaling has emerged as a powerful paradigm for improving large language model (LLM) performance through iterative inference-time refinement. Yet, its effectiveness typically relies on two conditions: (1) a well-defined problem space (e.g., mathematical reasoning), and (2) the model's ability to self-correct -- conditions that are rarely met in conversational product search. In this setting, user queries are often ambiguous and evolving, and MLLMs alone have difficulty grounding responses in a fixed product corpus. Motivated by these challenges, we propose a novel framework that introduces test-time scaling into conversational multimodal product retrieval. Our approach builds on a generative retriever, further augmented with a test-time reranking (TTR) mechanism that improves retrieval accuracy and better aligns results with evolving user intent throughout the dialogue. Experiments across multiple benchmarks show consistent improvements, with average gains of 14.5 points in MRR and 10.6 points in nDCG@1.

多模态对话推荐生成检索测试时扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。