arXiv:2606.23767cs.LG2026-06

用统一标准重评因果推断方法,发现零参数压缩基线表现最佳。

One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline

  • 所有方法在相同102对数据上统一运行,禁止调参并强制决策
  • 零参数压缩基线达74.7%准确率,超越多数论文报告结果
  • 揭露发表结果被高估的根源:测试集选模与选择性弃权

Tuebingen因果对的性能常被不同论文间比较,但各研究采用各自协议(子集、权重、模型选择、决策率)导致不可比。本文提出同一手重评估:我们统一运行所有方法于相同的102对数据,严格禁用调参且强制每对决策。引入一个零参数基线——排序-条件压缩:将量化、排序、一阶差分后的数据输入现成压缩器(bz2)。在统一标准下,排名显著不同于文献。该基线达74.7%加权准确率(p=3.7e-7);在与SLOPE同组的100对上得76.0%,仅比作者报告的强制决策版本低1.2点(77.2%),差异不显著(McNemar p=0.39)。RECI的复现结果为70.7%,落在原作者误差范围内,而非常被引用的77.5%(此误源于单元格复制错误)。SLOPE的82.4%是子集选择结果:仅在显著性检验选出的对上评分可复现为81.7%。在统一标准下,各方法集中于70%~75%区间,零参数压缩基线与最强方法持平。本文揭示了发布结果被高估的机制(测试集选模、显著性筛选弃权),并贡献两点新发现:压缩分数大小是模型无关的混淆标志(p=2.8e-68),预注册假验证失败,反向限定了方法理论解释范围。代码、预注册及每对输出均已公开。

原文摘要 · Abstract (English)

Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors' own protocol -- different pair subsets, weightings, model-selection, and decision rates. We argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one strict rule -- no tuning and a decision forced on every pair. As a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data to an off-the-shelf compressor (bz2) and has zero fitted parameters. Under the common ruler the ranking differs sharply from the literature. Our baseline reaches 74.7% weighted accuracy (p = 3.7e-7); on the same 100 pairs that SLOPE is evaluated on it scores 76.0%, a 1.2-point gap below the authors' own forced-decision SLOPE (77.2%) that is well inside noise (McNemar p = 0.39). A faithful re-run of RECI lands at 70.7% -- inside the original authors' reported error bar, not the 77.5% often quoted (which we trace to a mis-copied cell). SLOPE's published 82.4% is a decided-subset figure: scoring the authors' own stored output only on the pairs its significance test chose to answer reproduces 81.7%. Under the common ruler the methods cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them. We document the mechanisms that inflate published figures (test-set model selection, significance-gated abstention) and contribute two further results: compression score magnitude is a model-free confounding flag (p = 2.8e-68), and a pre-registered falsification test fails in an instructive way that bounds the method's theoretical interpretation. Code, pre-registrations, and per-pair outputs are released.

因果推断基准测试零参数可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。