arXiv:2608.21476cs.SEcs.CV2026-08

提出可审计的网站冗余评估框架,区分冗余的三种作用并保留完整证据链。

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

论文配图:From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy
图 1 · 摘自论文原文
  • 将冗余拆解为重复负载、正常使用负担和故障恢复储备三维度分别测量
  • 在测试中比单一指标方法更准确预测任务失败后的恢复能力
  • 保留原始截图与日志,支持人工复核,适合需透明验证的系统评测

网站冗余并无固定定义,相同重复元素在不同任务中可能造成干扰或提供备份。本文提出COR A(反事实、可观测冗余审计)框架,分别度量重复负载、正常使用负担与故障域恢复储备。每次运行均保留截图、稳定元素标识及任务轨迹。采用版本化视觉语言模型生成标注,经类型化验证与发布检查后决定是否报告校准维度;未通过或格式错误的输出保留在不可报告分母中。在透明机制测试平台上,因子化CORA表示能有效分离恢复储备与正常负担,并在扰动场景下预测任务成功率优于标量负载基线。模型研究显示,重复性不足保证有效性:两个小型本地视觉语言模型虽产生重复输出,但均未满足全部发布要求。CORA因此拒绝自动评分,仅保留原始响应与故障记录。独立检查模块确认了类型验证器与加固发布门禁符合设计规范;但这些测试无法证明其在生产站点上的语义基础或准确性。总体而言,CORA在此受控基准中被视为可审计的评估程序,而非通用标准。人类一致性、AI与人类判断的准确性,以及在独立生产站点上的验证仍属开放实证问题。

原文摘要 · Abstract (English)

Website redundancy does not have a single fixed meaning. The same repeated element may distract during one task and provide backup during another. We introduce CORA (Counterfactual, Observable Redundancy Audit), which measures repetition load, normal-use tax, and failure-domain recovery reserve separately. Each run retains screenshots, stable element identities, and task traces. A versioned vision-language model proposes the annotations. Typed validation and release checks then determine whether a calibrated dimension can be reported; failed or malformed outputs stay in the fixed denominator. On a transparent mechanistic testbed, the factorized CORA representation separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines. The model studies then showed why repeatability is not enough: two small local vision-language models produced recurring outputs, but neither instrument met all release requirements. CORA therefore withheld automated scores from both instruments while retaining the raw responses and failure records. Separate checker fixtures confirmed that the typed validator and hardened release gates implement their specifications; these tests do not establish semantic grounding or accuracy on production sites. Taken together, the results position CORA as an auditable candidate procedure for the controlled benchmark studied here rather than a general standard. Human agreement, AI-versus-human accuracy, and validation on independent production sites remain open empirical questions.

可审计评估网站冗余视觉语言模型测试验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。