让推荐系统像人一样看屏幕读文字,更懂用户对话意图。
Integrating Vision-Centric Text Understanding for Conversational Recommender Systems
- 用视觉化文本编码模拟快速扫屏,结合大模型精读关键信息。
- 在两个数据集上推荐准确率和回复质量均显著提升。
- 适合做智能对话推荐、多模态交互系统的研究者参考。
对话式推荐系统(CRS)通过自然语言交互提供个性化推荐,受到广泛关注。为更准确地从多轮对话中推断用户偏好,近期研究倾向于扩展对话上下文(如引入多样实体信息或检索相关对话)。然而,这种上下文丰富化会带来输入过长、文本风格不一致和无关噪声等问题,对语言理解能力提出更高要求。本文提出STARCRS——一种面向屏幕与文本感知的对话推荐系统,融合两种互补的文本理解模式:(1) 屏幕阅读路径,将辅助文本信息编码为视觉标记,模拟屏幕快速浏览;(2) 基于大模型的文本路径,聚焦有限关键内容进行细粒度推理。设计了基于知识对齐的融合框架,结合对比对齐、交叉注意力交互与自适应门控机制,实现双模式协同,提升偏好建模与响应生成效果。在两个常用基准数据集上的大量实验表明,STARCRS在推荐准确率和生成回复质量上均有持续提升。
原文摘要 · Abstract (English)
Conversational Recommender Systems (CRSs) have attracted growing attention for their ability to deliver personalized recommendations through natural language interactions. To more accurately infer user preferences from multi-turn conversations, recent works increasingly expand conversational context (e.g., by incorporating diverse entity information or retrieving related dialogues). While such context enrichment can assist preference modeling, it also introduces longer and more heterogeneous inputs, leading to practical issues such as input length constraints, text style inconsistency, and irrelevant textual noise, thereby raising the demand for stronger language understanding ability. In this paper, we propose STARCRS, a Screen-Text-AwaRe Conversational Recommender System that integrates two complementary text understanding modes: (1) a screen-reading pathway that encodes auxiliary textual information as visual tokens, mimicking skim reading on a screen, and (2) an LLM-based textual pathway that focuses on a limited set of critical content for fine-grained reasoning. We design a knowledge-anchored fusion framework that combines contrastive alignment, cross-attention interaction, and adaptive gating to integrate the two modes for improved preference modeling and response generation. Extensive experiments on two widely used benchmarks demonstrate that STARCRS consistently improves both recommendation accuracy and generated response quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。