arXiv:2606.30412cs.CYcs.AI2026-06中稿 · AAAI

用三角循环和多轮测试评估大模型排序可靠性,避免误判高风险决策。

Can LLMs Rank? A Tale of Triads and Triage

论文配图:Can LLMs Rank? A Tale of Triads and Triage
图 1 · 摘自论文原文
  • 通过检测成对比较中的矛盾三角形,量化模型内部一致性。
  • 不同大模型在两项高风险任务中表现差异显著,可靠性不一。
  • 建议同时使用两种指标,确保排序结果可信可用。

从无家可归者住房分配到急诊科分诊,大语言模型正被用于对重要资源进行人员排序。同时排序大规模群体认知负荷高且易出错。借鉴社会选择理论,可通过成对比较并聚合为总序来解决。但关键问题在于:如何在投入排名前判断大模型的判断是否足够一致?本文提出两种一致性检测方法:经典系数ζ通过统计锦标赛图中的循环三角形,提供低成本、无需模型的运行内一致性度量;而肯德尔τ等距离度量则评估跨轮次的排序变异。理论与实践均表明二者独立有效,应联合使用以评估排序可靠性。我们在无家可归者服务分配与急诊分诊两个高风险任务中验证了其重要性,三种主流大模型在一致性维度上表现迥异。本文为从业者提供了排序前评估一致性的实用指南。

原文摘要 · Abstract (English)

From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources. Ranking large groups simultaneously is cognitively demanding and error-prone. A natural solution, drawing on decades of social choice theory, elicits pairwise comparisons and aggregates them into a total order. However, a fundamental question remains when LLMs serve as the pairwise judge: how can a practitioner tell, before committing to a ranking, whether the LLM's judgments are sufficiently consistent to trust the result? We discuss two different ways of identifying consistency. A classical diagnostic, the coefficient of consistency $ζ$, originally developed to measure judge reliability by counting circular triads in tournament graphs, provides a cheap, model-free measure of intra-run consistency. Various standard measures of distance between rankings, for example Kendall's $τ$, can measure inter-run variability. We show, in both theory and practice, that these measures are independently valuable, and advocate for using both to assess reliability of rankings. We demonstrate the practical importance of our results across two high-stakes prioritization tasks: homelessness service allocation and emergency department triage. Three different leading LLMs have considerably different performance profiles across these two axes of consistency. We provide guidelines for how practitioners could think about measuring and assessing consistency before committing to a model for ranking or prioritization.

大模型排序可靠性评估社会选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。