arXiv:2602.09829cs.IR2026-02ACL被引 5

将多智能体推理内化为单模型,实现高效精准推荐

Internalizing Multi-Agent Reasoning for Accurate and Efficient LLM-based Recommendation

  • 用多智能体教师生成带反思的推理过程
  • 蒸馏后模型比教师快39.5%且精度更高
  • 适合需要实时推理的推荐系统场景

大型语言模型通过利用广泛的世界知识和语义推理能力,正在重塑推荐系统以理解用户意图。然而,如何有效融合协同信号并避免高昂的推理延迟仍是关键瓶颈。为此,我们提出一种轨迹驱动的内化框架,构建单智能体轨迹对齐推荐器(STAR)。首先设计一个多智能体教师系统,支持多轮工具使用与自我反思;该教师采用协同信号转换机制,将隐含行为模式转化为自然语言描述证据,提升推理准确性。随后,通过轨迹驱动的蒸馏流程,将包括规划、工具使用和自省在内的代理逻辑转移至紧凑的STAR模型中。大量实验表明,STAR在保持无迭代延迟的前提下,相比其教师模型性能提升8.7%至39.5%,为实时、增强推理的推荐系统铺平了道路。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are reshaping recommender systems by leveraging extensive world knowledge and semantic reasoning to interpret user intent. However, effectively integrating these capabilities with collaborative signals while avoiding prohibitive inference latency remains a critical bottleneck. To address this, we propose a trajectory-driven internalization framework to develop a Single-agent Trajectory-Aligned Recommender (STAR). Specifically, to internalize complex reasoning capabilities into a single efficient model, we first design a multi-agent teacher system capable of multi-turn tool usage and reflection. This teacher utilizes a Collaborative Signal Translation mechanism to explicitly convert latent behavioral patterns into descriptive natural language evidence to enhance reasoning accuracy. Subsequently, a trajectory-driven distillation pipeline transfers this agentic logic, including planning, tool usage, and self-reflection, into the compact STAR model. Extensive experiments demonstrate that STAR surpasses its teacher by 8.7% to 39.5% while eliminating iterative latency, paving the way for real-time, reasoning-enhanced recommendation.

推荐系统多智能体推理内化效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。