arXiv:2605.18805cs.IRcs.AI2026-05

评测大模型推荐代理的实用价值,不只看说得对不对,更看有没有真用。

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents

  • 用真实用户行为数据训练推荐效用代理,评估推荐集的相关性、互补性和多样性。
  • 在控制环境下测试发现,模型表现随算力和工具质量提升,但语义合理不等于实际有用。
  • 适合研发购物助手的人,想让推荐既靠谱又真正帮上忙。

大语言模型推荐代理生成带有自然语言解释的物品集合报告。现有评估多简化为小范围重排序或仅关注语义合理性。我们提出推荐图谱(RecoAtlas),一个基于行为指标的购物代理评测基准与工具包。RecoAtlas结合保留交互数据的效用代理,评估相关性、互补性和多样性,同时独立衡量语义连贯性与解释质量。其受控工具环境使代理面对语义型、行为对齐型或故障型工具,可诊断性能提升源于更强推理、更好信号还是更优工具使用策略。实验表明,RecoAtlas具备有意义基准的关键特性:性能随模型容量和测试时计算资源增长,依赖强且对齐的工具,受噪声或错位信号影响而下降,并揭示语义合理性不等同于行为驱动效用。RecoAtlas为开发优化推荐集实用性、一致性与行为依据的购物助手奠定基础。

原文摘要 · Abstract (English)

LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet existing evaluations often reduce this setting to reranking small shortlisted candidate sets or judge reports mainly by semantic plausibility. We introduce Recommendation Atlas (Agentic Tool-Level Assessment for Shopping), or RecoAtlas, a benchmark and toolkit for evaluating shopping agents with behavior-grounded metrics. RecoAtlas complements held-out interaction metrics with learned utility proxies for relevance, complementarity, and diversity derived from interaction data, while separately measuring semantic coherence and explanation quality. Its controlled tool environment exposes agents to either semantic, behavior-aligned, or faulty tools, enabling diagnosis of whether performance gains arise from stronger reasoning, better signals, or more effective tool-use policies. Across controlled experiments, we show that RecoAtlas exhibits key properties of a meaningful benchmark for agentic systems: performance scales with model capacity and test-time compute, improves with stronger and better-aligned tools, degrades under noisy or misaligned signals, and reveals that semantic plausibility does not necessarily capture behavior-grounded utility. RecoAtlas provides a foundation for developing and evaluating shopping assistants that optimize not only for plausible recommendations, but also for coherent, behaviorally grounded recommendation sets.

推荐系统大模型应用行为评估工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。