用机器学习采样+大模型标注,精准测量平台违规内容的真实曝光率。
Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling
- 基于机器学习加权采样,聚焦高曝光高风险内容,节省人工成本。
- 每日采集样本可支持多维度分析,误差可控,结果带置信区间。
- 适合内容安全团队做跨平台、跨区域的违规内容监测与决策参考。
内容安全团队需要反映用户真实体验的指标,而非仅依赖报告数据。本文研究‘流行度’:某日用户观看中违反特定政策的内容占比。准确测量流行度困难,因违规内容稀少且人工标注成本高,难以频繁开展具有平台代表性的研究。我们提出一种基于设计的测量系统:(i)利用机器学习辅助权重,从曝光流中每日抽取概率样本,将标注预算集中于高曝光、高风险内容,同时保持无偏性;(ii)使用多模态大模型结合政策提示和黄金标注集进行标签生成;(iii)产出符合设计的一致性流行度估计,包含置信区间与仪表盘下钻分析。核心设计目标是单一全局样本支持多维度分析:同一日样本可通过事后分层估计,支持按平台、用户地理、内容年龄等细分维度的流行度测算。本文详述统计估计器、方差与置信区间构建方法、标签质量监控机制,以及可配置的工程工作流,适用于多种政策场景。
原文摘要 · Abstract (English)
Content safety teams need metrics that reflect what users actually experience, not only what is reported. We study prevalence: the fraction of user views (impressions) that went to content violating a given policy on a given day. Accurate prevalence measurement is challenging because violations are often rare and human labeling is costly, making frequent, platform-representative studies slow. We present a design-based measurement system that (i) draws daily probability samples from the impression stream using ML-assisted weights to concentrate label budget on high-exposure and high-risk content while preserving unbiasedness, (ii) labels sampled items with a multimodal LLM governed by policy prompts and gold-set validation, and (iii) produces design-consistent prevalence estimates with confidence intervals and dashboard drilldowns. A key design goal is one global sample with many pivots: the same daily sample supports prevalence by surface, viewer geography, content age, and other segments through post-stratified estimation. We describe the statistical estimators, variance and confidence interval construction, label-quality monitoring, and an engineering workflow that makes the system configurable across policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。