arXiv:2608.18300cs.AI2026-08

用生命周期管理大模型评审员,提升推荐解释的实效性。

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

论文配图:The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
图 1 · 摘自论文原文
  • 构建评审员全周期框架,涵盖标准定义、训练调优、部署评估与持续监控。
  • 五周实验显示,经评审员优化的解释使用户观看新内容比例上升,浏览转播放率提高。
  • 适合关注推荐系统可解释性与大模型持续运维的研究者和工程师。

LLM-as-a-Judge 通过大语言模型评估另一AI生成的自然语言,已成为加速和扩展昂贵人工评估的标准方法。然而多数研究将评审器视为静态产物,仅在构建时或固定基准上评估一次。本文提出,部署中的LLM评审器应被视为具有生命周期的动态系统,需经历构建、训练、部署与持续维护四个阶段,每个阶段均面临独特技术与运营挑战。我们以Netflix推荐解释的评审系统为例,基于一系列受控在线实验,每周生成并评估数十万条不同剧集级别的解释,覆盖不断变化的内容库。框架包含四阶段:(I) Birth 定义评估标准,构建含人类标注与推理依据的精选数据集;(II) Training 通过推理对齐评分调优(RART),利用元评审器对推理输出进行学习信号反馈;(III) Deployment 将评审器投入质量拦截与反思式生成双重角色;(IV) Monitoring 实施持续的人机协同(HITL)对齐机制,检测偏差并触发人类审核后的再调优。五周在线A/B测试覆盖数千万用户,结果显示,经评审器对齐的解释使用户更倾向观看未看过的新型内容,并提升浏览转播放成功率,且无质量相关问题升级。

原文摘要 · Abstract (English)

LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. Yet most work treats a judge as a static artifact, evaluating it once at construction or against a fixed benchmark. We argue instead that an LLM judge operating in a deployed system is better understood as having a lifecycle. It must be built, trained, deployed, and continuously maintained as the surrounding data evolves, and each phase poses distinct technical and operational challenges. We present such a lifecycle for the LLM judges that evaluate recommendation explanations at Netflix. Everything we report comes out of a series of controlled online member-facing experiments, in which our pipeline generated and the judges assessed hundreds of thousands of distinct show-level explanations per week across a changing catalog. Our framework has four phases. (I) Birth defines the evaluation criteria and builds curated benchmark datasets with human labels and rationales. (II) Training refines the judges' rubrics via Reasoning-Aligned Rubric Tuning (RART), which uses a meta-judge over reasoning output as the learning signal. (III) Deployment puts one judge in two online roles, quality gating and reflective generation. (IV) Monitoring runs a continuous Human-in-the-Loop (HITL) alignment process that detects drift and triggers re-tuning behind a human review gate. We report results from a five-week online A/B test over tens of millions of members on the Netflix mobile app, in which judge-aligned explanations shifted member viewing toward novel content (previously unwatched) and increased successful browse-to-play sessions relative to a no-explanation control, with no quality-related escalations.

推荐系统大模型评审生命周期管理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。