AI评估结果不能自动推广,需检查推理链条是否可靠
When benchmark inferences do not compose: Projectibility in AI evaluation
- 提出可操作的评估链审计方法,识别无效推断
- 发现不同研究间条件变化导致结论无法衔接
- 适合评估可信度与系统部署决策者使用
AI基准测试结果很少能一步得出有影响力的结论。评估者需将其推广至新场景、解释为能力证据、外推至新任务、迁移至其他系统或环境,并结合对人工评审和下游影响的假设。以有效性为中心的方法要求每一步都有证据支持。本文明确并操作化了此类方法留给分析者的难题:合理链接不等于合理链路。一项研究的目标未必是下一项的来源;系统、群体、结果或条件可能在接口处发生变化;共享数据或模型谱系可能使看似独立的支持变得依赖。项目性(Projectibility)关注的是从已观察到未观察案例的有限延伸是否合理。古德曼提出了竞争性外推问题;基于论证的有效性提供了检验框架。本文贡献是一种分布式AI证据的接口审计:定义类型化的源与目标描述,并建立程序,区分从未相遇的端点与虽相遇但验证未通过的端点。法律研究案例显示,基准证据与部署研究均可有效,却仍保持平行。已知真值演示表明,聚合稳定性可能掩盖后续投影所需的关键差异。由此产生的项目性审计可诊断评估到应用中的不成立连接。
原文摘要 · Abstract (English)
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper makes explicit and operationalizes a problem those approaches leave to the analyst: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The contribution is an interface audit for distributed AI evidence: typed source and target descriptions, and a procedure separating endpoints that never meet from endpoints that meet while warrant fails to cross. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A known-truth demonstration shows why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。