提出无需依赖昂贵模型的可靠内容评分方法,可检测并抵抗操纵。
Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End
- 用对抗性验证门控机制构建确定性评分代理,确保结果可信。
- 在500个对抗编辑样本上,攻击者最多只能提升6分且随剂量递减。
- 适合需要防篡改评估的生成模型部署与质量过滤场景。
如何验证一个廉价、确定性的代理评分,来替代昂贵、限速且不稳定的权威评估?我们提出一套基于对抗性伪造检验门控(负向控制、剂量响应、有限放大、重复惩罚、长度中性)的协议,定义并筛选出该代理,在训练集上拟合并在保留集上验证;围绕这些门控,界定代理无法解决的边界,并对当前权威模型重新测量外部因果证据而非假设其有效性。我们在生成引擎优化任务中端到端验证该方法:代理为确定性内容评分,结果显示唯一已知的因果锚点(2023年效应量)在十个现代引擎家族中均未显著影响引用数,说明锚点已失效;重新校准至近零现代向量后,评分移除了所有杠杆响应成分。剩余部分为门控强制的响应曲面。在500源对抗编辑基准测试中,放大校准后的杠杆最多仅使攻击者得分提升6分,且随剂量递减;单杠杆放大有理论上限,而跨杠杆亚加性为实证发现,与之一致。检测时,网络垃圾基线占主导,分布外攻击可规避评分,因此在其上部署可部署的过滤层。查询条件下的最优前沿(skyline)限制了评分的引用信号(组内斯皮尔曼相关系数0.11),将原本与查询无关的评分定位为质量过滤器而非引用预测器。首次排名评估中的查询泄露漏洞和信心标志失败均已披露并修正,所有数据均可从发布资源离线复现,边际API成本为零。
原文摘要 · Abstract (English)
How do you validate a cheap, deterministic proxy for an oracle that is expensive, rate-limited, and non-stationary? We present a protocol built on adversarial falsification gates (negative control, dose response, bounded amplification, duplication penalty, length neutrality) that define and select the proxy, fitted on a training split and confirmed held-out; around them it bounds what the proxy can never resolve, and re-measures external causal evidence on the current oracle rather than assuming it. We demonstrate it end to end on Generative Engine Optimization, where the proxy is a deterministic content score, and one step fails on that domain exactly as the protocol is built to detect: re-measuring the only published causal anchors (2023 effect sizes) on ten modern engine families shows their levers move citation on none, so the anchors are an expired external check; recalibrating to the near-zero modern vector strips the score of its lever-responsive components. What survives is the gate-enforced response surface. The gates buy a measured property: on a 500-source benchmark of adversarial edits, amplifying the score's calibrated levers gains an attacker at most 6 points, and decreases with dose; single-lever amplification is provably bounded, while the cap and cross-lever sub-additivity are empirical findings consistent with it. On detection, web-spam baselines dominate and out-of-distribution attacks evade the score, so the deployable filter layers it over them. A query-conditioned skyline bounds the score's citation signal (within-query Spearman 0.11), repositioning query-agnostic scores as quality filters rather than citation predictors. A query-leakage bug in our first ranking evaluation and a failed confidence flag are disclosed and corrected; every number reproduces offline from released artifacts at zero marginal API cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。