arXiv:2608.23154cs.IR2026-08中稿 · the Recsys'26 Work…

提升推理过程可解释性,反而降低推荐效果。

The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness

  • 用自然语言标题与语义标识符分别生成推理链,对比可解释性。
  • 即使推理链更清晰,推荐准确率仍下降,且无法通过强化学习完全恢复。
  • 适合关注推荐系统可解释性与性能权衡的研究者。

近期研究致力于提升生成式推荐中的显式自然语言推理链质量,包括在语义标识符(SID)预测中引入思维链推理。然而,由于SID是难以理解的模型学习标识符,需经过大量对齐才能让大模型进行推理,这为实验提供了可控条件:可独立改变物品表示方式(标题 vs. SID)和语义对齐程度(最小 vs. 全面)。我们在三个亚马逊商品领域,基于统一的Qwen3-1.7B模型,进行了2×2因子实验。结果表明,引入显式推理链反而降低了标准SFT和强化学习训练下的离线推荐效果,尽管自然语言标题生成的推理链更具语义可解释性。全面的SID对齐虽提升了推理链质量,但未改善推荐效果,仅更丰富的奖励信号部分恢复了性能。总体而言,提升推理链质量本身不足以持续提高传统离线推荐效果,尤其是在当前训练目标与评估协议下。

原文摘要 · Abstract (English)

Recent work has focused on improving explicit natural-language descriptive reasoning traces for generative recommendation. This includes systems that augment semantic ID (SID) prediction with chain-of-thought reasoning. However, because SIDs are opaque learned identifiers rather than natural language, they require costly alignment before an LLM can reason over them. This provides a controlled experimental setting in which both item representation (Title vs. SID) and semantic grounding (minimal vs. extensive SID alignment) can be varied independently. We therefore present the first controlled comparison of descriptive reasoning trace quality across semantic IDs and natural-language titles in a 2 x 2 factorial study on three Amazon product domains using a shared Qwen3-1.7B backbone. We find that introducing explicit descriptive reasoning traces reduces traditional offline recommendation effectiveness under standard SFT and RL training, even though natural language titles produce substantially more grounded and interpretable traces. Extensive SID alignment improves descriptive trace quality but not traditional offline recommendation effectiveness, while a richer reward signal partially recovers performance. Overall, our results show that improving descriptive reasoning trace quality is not, by itself, sufficient to consistently improve traditional offline recommendation effectiveness under the training objectives and evaluation protocols studied here.

推荐系统可解释性大模型推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。