为评估系统设计可信的投票机制,确保结果可复现。
Aggregate Disambiguation Systems

- 用有限样本估计评审团与总体决策一致的概率。
- 给出在容忍度内分歧概率不超限的解的最低占比。
- 适用于需保证评估可复现性的评测系统设计者。
自然语言任务可能引发遵循协议的评审员对相同信息得出不同结论。本文研究聚合消歧系统(ADS)。给定任务和候选方案,每位评审员以二元投票决定是否接受该方案,系统聚合有限评审团的投票结果。目标是相对于显式声明的评审员参考群体实现协议可复现性,而非语义真实性。我们区分固定有限普查、概率性评审员群体和增长型普查极限,因它们的终点律与保证不可互换。在群体设定下,利用有限样本估计有限评审团与声明的评审员群体达成相同决策的频率。提供一个下置信界,用于估计在指定容差内分歧概率不超过的候选方案比例。计算分别考虑候选方案采样与评审员采样。构造允许列间任意依赖(由共享评审员行引起),在评审员层使用精确二项区间,在生成器层采用精确单侧二项反演。模拟验证了实现效果,并揭示了统计功效的局限性。
原文摘要 · Abstract (English)
Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregate disambiguation systems (ADSs). Given a task and a candidate solution, each evaluator casts a binary vote on whether the solution should be accepted, and the system aggregates the votes of a finite panel. The target is protocol reproducibility relative to an explicitly declared evaluator reference, not semantic truth. We separate fixed finite censuses, probabilistic evaluator populations, and growing-census limits, since their endpoint laws and guarantees are not interchangeable. In the population setting, we use finite samples to estimate how often a finite panel reaches the same decision as the declared evaluator population. We provide a lower confidence bound on the fraction of candidate solutions for which the disagreement probability is at most a chosen tolerance. The calculation accounts separately for sampling candidate solutions and sampling evaluators. The construction permits arbitrary dependence among columns induced by shared evaluator rows and uses exact binomial intervals at the evaluator layer and an exact one-sided binomial inversion at the generator layer. Simulations check the implementation against known population coverages and expose power limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。