提出防御平台操纵的安全审计方法,确保评分真实反映危害降低。
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
- 用语义等价类构建转换图,识别可被操纵的评分漏洞。
- 语义包络评分法在所有测试中无操纵漏洞,优于传统指标。
- 适合关注合规审计可靠性的政策制定者与平台安全团队。
英国《在线安全法》和欧盟《数字服务法》日益将量化指标作为合规证据。一旦公布,平台便可能通过路由推荐至语义等价的内容变体来提升分数,而无需真正减少危害。本文研究此类审计指标何时仍能证明危害的真实下降。将协议建模为公开的变换图,其连通分量构成语义类,指标本身视为安全对象。三个结论:第一,任何直接对变体评分的指标,只要同一有害类内两个等价变体得分不同,即存在可操纵性;第二,语义包络提升(semantic-envelope lift)是保守类常数修复中点态最小的唯一方案;第三,对所有平台策略成立的分层证书 $H^/star(x) \le (1/\hatα) M_{\mathrm{Env}(m)}(x) + \barη$,其中 $\barη$ 吸收标注与协议误差。通过三种方式验证:有限状态混合策略网格的穷举枚举、Z3 与 cvc5 的 SMT 编码交叉回放、以及 PRISM-games 编码的有界单玩家马尔可夫决策过程。脆弱指标在操纵不变性上失败,无法支持预声明的类覆盖证书;在包络级证书下,所有测试实例均出现显著违规,且在固定审计预算下随机目录的平均操纵差距巨大。语义包络指标在所有测试实例中均无违规。
原文摘要 · Abstract (English)
Online-safety regulation under the UK Online Safety Act and the EU Digital Services Act increasingly treats scalar metrics as compliance evidence. Once announced, such a metric also becomes an optimization target: a strategic platform can improve its score by routing recommendations through semantically equivalent content variants, without reducing true harm. We ask when such an audit metric can still certify a genuine reduction in harm. The protocol is modeled as a published transformation graph whose connected components form semantic classes, and the metric itself is treated as a security object. Three results follow. First, any metric that scores variants directly is manipulable as soon as two equivalent variants in a harmful class disagree in score. Second, the semantic-envelope lift, which assigns each variant the maximum score in its class, is the unique pointwise minimum among conservative classwise-constant repairs. Third, a class-stratified certificate, $H^\star(x) \le (1/\hatα) M_{\mathrm{Env}(m)}(x) + \barη$, holds for every platform strategy, with $\barη$ absorbing annotation and protocol error. We check the claims at three levels: exhaustive enumeration on a finite-state grid of mixed strategies, an SMT encoding in Z3 cross-replayed in cvc5, and a bounded single-player MDP encoded in PRISM-games. The fragile metric fails manipulation invariance and cannot support the same useful predeclared class-coverage certificate; under the envelope-level certificate, it produces large violations at every tested instance, with a large mean gaming gap across random catalogs at a fixed audit budget. The semantic-envelope metric exhibits no such violation in the tested instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。