arXiv:2508.00129cs.AImath.OC2025-08

首次实现多准则决策分析中排名反转的自动化检测,揭示其在文献中普遍存在。

Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis

  • 提出可落地的算法框架,将三类检测标准转为可执行流程。
  • 实证发现近半数文献存在重组一致性失效,14.8%违反传递性。
  • 开源工具支持真实场景应用,推动方法可靠性评估标准化。

排名反转指备选方案排序违背理性决策公理,是多准则决策分析(MCDA)方法可信度的重大威胁。Wang与Triantaphyllou(2008)提出三项系统测试标准,但17年来始终缺乏经验证的开源实现,阻碍理论向实践转化。本文在开源库Scikit-Criteria中构建算法框架,将三类标准转化为可集成于实际工作流的程序:基于分层消歧的受控降级策略(RRT1),以及精确凝聚与传递约简的支配图构造法(RRT2/RRT3)。通过两个案例验证:一是加密货币评估问题(Van Heerden et al., 2021),二是对20篇已发表论文中27个管道/数据组合的大规模审计。结果显示,顶级备选方案稳定性(RRT1)达96.3%,但传递性(RRT2)失效比例为14.8%,最严格的标准重组一致性(RRT3)失效率高达48.1%,表明排名反转是当前MCDA文献中的普遍、可量化的现象,而非偶发或人为攻击所致。

原文摘要 · Abstract (English)

Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-documented threat to the reliability of Multi-Criteria Decision Analysis (MCDA) methods. Wang and Triantaphyllou (2008) proposed three systematic test criteria to detect this phenomenon, but despite more than 700 citations, no validated, open-source implementation has closed the gap between theory and practice, a 17-year absence we trace to the non-trivial algorithmic challenges of operationalizing these tests for real-world pipelines. We present an algorithmic framework, implemented in the open-source Scikit-Criteria library, that translates Wang and Triantaphyllou (2008)'s three criteria into concrete, pipeline-compatible procedures: a controlled degradation strategy with hierarchical tie-breaking and graceful handling of preprocessing filters (RRT1), and a dominance-graph construction with exact condensation and transitive reduction for detecting transitivity violations and recomposition inconsistencies (RRT2/RRT3). We demonstrate the framework through two case studies: an application to a cryptocurrency evaluation problem (Van Heerden et al., 2021), and a large-scale audit of 27 pipeline/dataset combinations reproduced from 20 published MCDM methods. The audit shows top-alternative stability (RRT1) is nearly universal (96.3%), but transitivity (RRT2) fails for 14.8% and recomposition consistency (RRT3), the strictest criterion, fails for nearly half (48.1%) of published examples-evidence that rank reversal is a pervasive, measurable feature of the current MCDM literature, not a marginal or adversarial concern.

决策分析排名反转算法验证开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。