arXiv:2608.19140cs.AIcs.CY2026-08

评估AI模型应看输出稳定性,而非平均能力。

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

  • 用重复任务测试输出一致性,无需人工评分
  • 顶尖模型间差异在输出集中度,而非平均表现
  • 帮助区分可调参数问题与需换模型的根本缺陷

当前对前沿语言模型的比较和评测聚焦于能力——其最佳或平均输出能达到什么水平。我认为这衡量的是错误维度。模型的准确率已趋于饱和:平均输出已落在目标位置。如今实际区分系统差异的是精度:在多次相同请求下,输出围绕目标的聚集程度。借鉴射手的区分,能力是平均落点,而可靠性是弹着点的聚集范围。本文提出三点主张:第一,精度而非能力才是系统间真正的前沿分水岭,但评测文化系统性地忽略了这一点,仅报告中心趋势而忽视分布;第二,精度可低成本、无循环地测量:固定一组确定性评分任务,在固定温度下多次运行,计算每项任务结果的一致性,无需模型内评分器;第三,该测量不仅是描述性的,更是决策导向的:能区分一致失败(群体偏离中心,可通过调整操作规则修正)与分散失败(群体宽泛,只能通过更换模型或采样策略解决)。本文定义了聚类度量,设计了测评框架,并展示追踪人机协作组群随时间变化,可获得论文1所需的复合信号。一次真实实验表明,一个差距可通过单条规则完全弥补(0/5 → 5/5),而基于规则自创的任务套件却未发现价值,因为前沿模型已内化良好实践——证明某规范的价值必须通过真实工作中的测量来验证,而非由其自身规则书构建。

原文摘要 · Abstract (English)

Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.

AI评测模型精度人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。