构建可验证的推荐系统评估基准,解决对话型推荐中评价主观性问题。
$τ$-Rec: A Verifiable Benchmark for Agentic Recommender Systems
- 用可验证奖励与揭示标签机制控制任务约束暴露方式。
- 最佳模型在pass^1下仅达57%,pass^4时降至35%。
- 适合评估对话式推荐系统的推理一致性,尤其关注大模型部署可靠性。
随着推荐系统向智能体化、多轮对话界面演进,评估方法难以跟上步伐。现有基准多依赖‘大模型作为评判者’,带来主观性、高成本和不一致问题。我们提出τ-Rec,一个面向智能体推荐系统的基准,以可验证奖励替代主观评价,并引入揭示标签诱发(RTE)机制,控制任务约束在对话中的显现方式。通过测试代理对结构化目录谓词的响应,并采用pass^k可靠性指标,τ-Rec实现对一致推理的系统性检验。我们在五种模型族的九种配置上进行评估——GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash、DeepSeek V4 Flash、Qwen3-32B 和 GPT-5 mini——结果显示显著的可靠性悬崖:即使最优模型在pass^1下也仅达约57%,pass^4时降至约35%,暴露出当前对话智能体部署中的关键差距。所有代码与数据均公开于https://github.com/nbharaths/tau-rec。
原文摘要 · Abstract (English)
As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on "LLM-as-a-judge" evaluations, which introduce subjectivity, high costs and inconsistency. We present $τ$-Rec, a benchmark for agentic recommender systems that replaces subjective evaluation with verifiable rewards and a reveal-tagged elicitation (RTE) mechanism that controls how task constraints surface during dialogue. By testing agents against structured catalog predicates and employing a pass^k reliability metric, $τ$-Rec provides a systematic test for consistent reasoning. Our evaluation of nine configurations across five model families -- GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, DeepSeek V4 Flash, Qwen3-32B and GPT-5 mini -- reveals a steep reliability cliff, where even the best model achieves only ~57% at pass^1 and ~35% at pass^4, highlighting a critical gap in current conversational agent deployment. All code and data are publicly available at https://github.com/nbharaths/tau-rec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。