arXiv:2503.00501cs.IRcs.CL2025-03被引 13

Qilin数据集聚焦多模态搜索推荐,基于小红书用户会话构建。

Qilin: A Multimodal Information Retrieval Dataset with APP-level User Sessions

  • 从3亿月活平台收集真实用户会话,涵盖图文、视频等多模态结果
  • 包含用户偏好答案及引用结果,支持RAG模型训练与行为分析
  • 适合研究多模态检索、用户行为建模与搜索推荐系统优化的学者

用户生成内容(UGC)社区通过整合视觉与文本信息提升用户体验,但现有高质量数据集稀缺限制了多模态搜索与推荐(S&R)研究进展。为此,本文提出Qilin数据集,源自拥有超3亿月活跃用户的社交平台小红书,平均搜索渗透率超70%。该数据集提供包含图文笔记、视频笔记、商业笔记及直接回答的异构结果,支持多样任务场景下的先进多模态神经检索模型开发。同时,收集了丰富的应用级上下文信号与真实用户反馈,尤其包含触发深度查询问答(DQA)模块的搜索请求中用户偏好的答案及其引用结果。这不仅支持检索增强生成(RAG)管道的训练与评估,还可探究DQA模块对用户行为的影响。通过全面分析与实验,揭示了优化S&R系统的若干洞见。

原文摘要 · Abstract (English)

User-generated content (UGC) communities, especially those featuring multimodal content, improve user experiences by integrating visual and textual information into results (or items). The challenge of improving user experiences in complex systems with search and recommendation (S\&R) services has drawn significant attention from both academia and industry these years. However, the lack of high-quality datasets has limited the research progress on multimodal S\&R. To address the growing need for developing better S\&R services, we present a novel multimodal information retrieval dataset in this paper, namely Qilin. The dataset is collected from Xiaohongshu, a popular social platform with over 300 million monthly active users and an average search penetration rate of over 70\%. In contrast to existing datasets, \textsf{Qilin} offers a comprehensive collection of user sessions with heterogeneous results like image-text notes, video notes, commercial notes, and direct answers, facilitating the development of advanced multimodal neural retrieval models across diverse task settings. To better model user satisfaction and support the analysis of heterogeneous user behaviors, we also collect extensive APP-level contextual signals and genuine user feedback. Notably, Qilin contains user-favored answers and their referred results for search requests triggering the Deep Query Answering (DQA) module. This allows not only the training \& evaluation of a Retrieval-augmented Generation (RAG) pipeline, but also the exploration of how such a module would affect users' search behavior. Through comprehensive analysis and experiments, we provide interesting findings and insights for further improving S\&R systems. We hope that \textsf{Qilin} will significantly contribute to the advancement of multimodal content platforms with S\&R services in the future.

多模态检索用户行为小红书RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。