PinSieve通过智能筛选提升企业内容质量审核效率,实现更快更准的自动化处理。
PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage
- 仅处理上游模型未决的灰色区域,结合路由评分与人工介入控制流程
- 过滤效率提升2.05倍,审核生产力提高25.7%,成本降低16.2%
- 支持可追溯、可审计的闭环反馈机制,适合生产级企业内容治理场景
企业在生产环境中部署的AI代理需具备边界可控、状态可查、可观测和可管理的特点,而非完全自治。本文介绍大规模内容质量流水线中的实际案例PinSieve,其核心组件为选择性视觉语言模型(VLM)服务代理,仅在轻量级上游模型未能解决的灰色区域运行,实时暴露路由评分,并保留受控的人工升级通道。在此部分,系统比旧模块多过滤2.05倍的非操作项,同时轻微降低误漏率;上线后,审核效率提升25.7%,单位运营成本下降16.2%,信号交付从次日提速至当日。我们进一步提出受控记忆飞轮机制,以选择性反馈维护系统:被升级项默认审查,自动通过项主要通过审计采样标注。反馈记忆记录路由轨迹、观察路径、审计倾向及重播元数据,用于评估与调试。数据清洗代理采用受限的提议-验证循环,基于代表性、不确定性、时效性及新审重播,设置正例率与分数区间守卫后批量接受。在六个月内生产数据的链式月度刷新中,平均FNR@50%从代表随机重播下的17.73%降至13.29%。推理审查代理对教师生成的推理过程进行审计,支持保留/修复/删除决策。生产成效仅归因于部署的服务代理;重播与推理审查结果为离线或采样治理证据。同一服务代理方案已在多个内部信号任务中复用,显示其任务外延潜力。
原文摘要 · Abstract (English)
Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its deployed component is a selective vision-language-model (VLM) Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation. On this slice, the deployed system filters 2.05x more non-actionable items than the previous production module while slightly reducing estimated miss rate; after promotion, it improves review productivity by 25.7%, reduces normalized operating cost by 16.2%, and moves signal delivery from next-day to same-day. We then study maintenance through a governed memory flywheel under selective feedback, where escalated items are reviewed by default and auto-passed items are labeled mainly through audit sampling. Feedback Memory records routing traces, observation paths, audit propensities, and replay metadata for evaluation and debugging. The Data Curation Agent uses a bounded proposal-verifier loop over representative, uncertainty, recency, and fresh-review replay, with positive-rate and score-bin guardrails before batch acceptance. In chained monthly refresh over six months of production data, this design reduces average FNR@50% from 17.73% under representative random replay to 13.29%. A Reasoning Review Agent audits teacher-generated rationales and supports keep/repair/drop decisions. Production claims are attributed only to the deployed Serving Agent; replay and rationale-review results are offline or sampled-governance evidence. The same serving-agent recipe has been adopted to several additional internal signals, suggesting transferability beyond one task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。