arXiv:2603.17405cs.LG2026-03

为高维数据因果表示学习提供可复现的评估基准与综合评分

Causal Representation Learning on High-Dimensional Data: Benchmarks, Reproducibility, and Evaluation Metrics

  • 构建多维度评估体系,覆盖重建、解耦、因果发现等方向
  • 提出单一综合指标,统一衡量模型在多个任务中的表现
  • 系统审查代码复现性,推动领域规范发展

因果表示学习(CRL)模型旨在将高维数据转换到潜在空间,通过潜在变量间的因果关系实现干预生成反事实样本或修改现有数据。为促进模型开发与评估,学界已提出多种合成与真实数据集,各有优劣。实际应用中,CRL模型需在重建、解耦、因果发现和反事实推理等多个方向上表现稳健,并采用相应指标。然而,多方向评估使模型比较复杂,某模型可能在某一方向优异但在其他方向表现不佳。此外,复现性问题突出:公开代码与重复实验结果一致性至关重要。本研究批判性分析现有数据集的局限性,提出适合CRL模型开发的数据集应具备的关键特征;引入一个整合各评估方向的单一综合指标,为每种模型提供全面评分;最后回顾文献中的实现方案,评估其复现性,识别领域差距并提炼最佳实践。

原文摘要 · Abstract (English)

Causal representation learning (CRL) models aim to transform high-dimensional data into a latent space, enabling interventions to generate counterfactual samples or modify existing data based on the causal relationships among latent variables. To facilitate the development and evaluation of these models, a variety of synthetic and real-world datasets have been proposed, each with distinct advantages and limitations. For practical applications, CRL models must perform robustly across multiple evaluation directions, including reconstruction, disentanglement, causal discovery, and counterfactual reasoning, using appropriate metrics for each direction. However, this multi-directional evaluation can complicate model comparison, as a model may excel in some direction while under-performing in others. Another significant challenge in this field is reproducibility: the source code corresponding to published results must be publicly available, and repeated runs should yield performance consistent with the original reports. In this study, we critically analyzed the synthetic and real-world datasets currently employed in the literature, highlighting their limitations and proposing a set of essential characteristics for suitable datasets in CRL model development. We also introduce a single aggregate metric that consolidates performance across all evaluation directions, providing a comprehensive score for each model. Finally, we reviewed existing implementations from the literature and assessed them in terms of reproducibility, identifying gaps and best practices in the field.

因果学习表示学习评估基准可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。