用可扩展框架提升大模型搜索的评估质量,实现近人类判断的自动化。
SAGE: Scalable AI Governance & Evaluation

- 构建双向校准循环,融合政策、判例与大模型裁判,生成多维可执行评估标准。
- 通过教师-学生蒸馏将高精度判断压缩至原成本的1/92,支持工业级部署。
- 在领英搜索中实测提升0.25%日活用户,发现传统指标无法捕捉的模型退化问题。
大规模搜索系统中的相关性评估受制于人力监管能力与生产系统高吞吐需求之间的治理鸿沟。传统方法依赖点击率等代理指标或稀疏人工评审,常无法覆盖关键相关性失效。本文提出SAGE(可扩展人工智能治理与评估)框架,将高质量人工产品判断转化为可扩展的评估信号。SAGE核心为双向校准环路,自然语言政策、精选判例与大语言模型代理裁判协同演化,系统化解语义模糊与对齐偏差,将主观相关性判断转化为可执行的多维评分体系,达成接近人类水平的一致性。为弥合前沿模型推理与工业级推理间的差距,采用教师-学生蒸馏技术,将高保真判断压缩至原成本的92倍低开销的紧凑学生模型。该框架已在领英搜索生态中部署,通过模拟驱动开发指导模型迭代,提炼出符合政策的在线服务模型,并实现快速离线评估。实际应用中,支撑了政策监督,量化评估模型变体并检测到传统参与度指标无法发现的性能退化。整体推动领英日活跃用户提升0.25%。
原文摘要 · Abstract (English)
Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems. While traditional approaches rely on engagement proxies or sparse manual review, these methods often fail to capture the full scope of high-impact relevance failures. We present \textbf{SAGE} (Scalable AI Governance \& Evaluation), a framework that operationalizes high-quality human product judgment as a scalable evaluation signal. At the core of SAGE is a bidirectional calibration loop where natural-language \emph{Policy}, curated \emph{Precedent}, and an \emph{LLM Surrogate Judge} co-evolve. SAGE systematically resolves semantic ambiguities and misalignments, transforming subjective relevance judgment into an executable, multi-dimensional rubric with near human-level agreement. To bridge the gap between frontier model reasoning and industrial-scale inference, we apply teacher-student distillation to transfer high-fidelity judgments into compact student surrogates at \textbf{92$\times$} lower cost. Deployed within LinkedIn Search ecosystems, SAGE guided model iteration through simulation-driven development, distilling policy-aligned models for online serving and enabling rapid offline evaluation. In production, it powered policy oversight that measured ramped model variants and detected regressions invisible to engagement metrics. Collectively, these drove a \textbf{0.25\%} lift in LinkedIn daily active users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。