Pinterest构建评估内容审核决策质量的量化框架,提升AI与人工判断的可信度。
Decision Quality Evaluation Framework at Pinterest
- 用专家标注的黄金数据集作为基准,建立可信赖评估标准。
- 通过倾向得分采样自动扩展数据覆盖,提升评估效率。
- 适用于LLM性能对比、提示词优化和政策演变追踪,适合平台安全团队使用。
在线平台需在大规模下实施内容安全策略。审核决策的质量评估是关键环节,但受成本、规模与可信度之间的权衡及政策复杂性影响,难度较大。为此,我们提出并部署了Pinterest的决策质量评估框架。该框架以领域专家(SMEs)构建的高可信黄金集(GDS)为基准,引入基于倾向得分的自动化智能采样管道,高效扩展数据覆盖范围。框架在多个关键场景中验证:评估不同LLM代理的成本-性能权衡、实现数据驱动的提示词优化方法、管理复杂政策演化过程,并通过持续验证确保政策内容占比指标的完整性。该框架推动内容安全系统从主观判断转向数据驱动的量化管理。
原文摘要 · Abstract (English)
Online platforms require robust systems to enforce content safety policies at scale. A critical component of these systems is the ability to evaluate the quality of moderation decisions made by both human agents and Large Language Models (LLMs). However, this evaluation is challenging due to the inherent trade-offs between cost, scale, and trustworthiness, along with the complexity of evolving policies. To address this, we present a comprehensive Decision Quality Evaluation Framework developed and deployed at Pinterest. The framework is centered on a high-trust Golden Set (GDS) curated by subject matter experts (SMEs), which serves as a ground truth benchmark. We introduce an automated intelligent sampling pipeline that uses propensity scores to efficiently expand dataset coverage. We demonstrate the framework's practical application in several key areas: benchmarking the cost-performance trade-offs of various LLM agents, establishing a rigorous methodology for data-driven prompt optimization, managing complex policy evolution, and ensuring the integrity of policy content prevalence metrics via continuous validation. The framework enables a shift from subjective assessments to a data-driven and quantitative practice for managing content safety systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。