arXiv:2604.25580cs.CL2026-04中稿 · EMNLP

研究指出依赖外部工具测毒性会出问题,呼吁学术界自建可掌控的评估体系。

Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure

论文配图:Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure
图 1 · 摘自论文原文
  • 主张研究领域应自建并自主管理评估基础设施,而非依赖外部工具
  • 发现依赖该工具导致结论夸大、结果随模型重训练波动、无法解释数据差异
  • 适合关注模型评估伦理、开放科学与研究可持续性的学者参考

Perspective API 将于 2026 年关闭,这移除了毒性检测的行业标准,暴露了研究人员对其不可控工具的依赖。基于此案例,我们提出:研究领域必须自建并自主治理测量基础设施,而非借用。通过对 241 篇使用或研究 Perspective 的论文分析,我们揭示了依赖该工具的代价:研究声称超出其实际支持范围、结果因模型无声重训练而改变、以及可测量但无法解释的偏差。这些缺陷在整个大语言模型生命周期中被放大——Perspective 提供标签、过滤语料库、评估训练系统,甚至奖励错误而非识别错误。为确保关闭后仍可研究,我们发布了来自 77 个数据集的 590 万条文本片段的 Perspective 分数。此外,我们提出十项领域应拥有的测量基础设施要求,并指出阻碍建设的不是技术能力,而是领域对基础设施工作的价值认知。

原文摘要 · Abstract (English)

Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool they did not control. Drawing on this case, we argue that a research field must build and govern its own measurement infrastructure rather than borrow it. Surveying 241 papers that use or study Perspective, we show what depending on it cost the research community: claims reaching past what the tool could support, results that shifted when its model was silently retrained, and disparities researchers could measure but not explain. These failures were amplified throughout the LLM lifecycle, where Perspective supplied the labels, filtered the corpora, and graded the systems trained on each, rewarding errors rather than catching them. To keep the instrument open to study after shutdown, we release Perspective scores for 5.9 million text snippets from 77 datasets. We further specify ten requirements for measurement infrastructure a field owns, and argue that what blocks such infrastructure is not technical capability but the value the field places on infrastructure work.

评估体系大模型研究伦理开放科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。