arXiv:2606.05689cs.LG2026-06

区分进化选择与静态选择,提出新模型揭示演化数据背后的因果机制

Causal Modeling of Selection in Evolution

  • 区分静态筛选与进化筛选两类选择机制,提出针对性建模方法
  • 在多代演化数据中成功识别出真实因果结构,避免传统方法的误判
  • 适用于免疫适应、抗药性等演化过程分析,适合生物与社会学研究者

理解数据中的潜在选择机制对因果发现至关重要。我们指出,常见叙事中的“选择”可分为两种形式:静态选择与进化选择。静态选择指一次性筛选过程,如调查志愿者偏差;进化选择则通过反复的繁殖适配度差异作用,观测数据是历史演化轨迹塑造的最新一代,如免疫适应、抗生素耐药性和社会规范形成。现有方法通常混淆两者,依赖相同的图形模型,该模型仅适用于静态场景,在演化条件下会失效并导致错误发现。为此,我们提出一种专用于刻画进化选择的新模型,并开发了从一个或多个环境/世代数据中识别该模型的完整有效方法。实验结果验证了该方法从数据中还原演化机制的能力。

原文摘要 · Abstract (English)

Understanding potential selection in data is crucial for causal discovery; we argue that "selection" in common narratives takes two forms, which we term static and evolutionary selection, respectively. Static selection refers to a one-shot filtering process where observed data consist of a subset of the population of interest, as in survey volunteer bias. Evolutionary selection, in contrast, operates through repeated rounds of differential fitness in reproduction, where observed data constitute the latest generation shaped by a historical trajectory, as in immune adaptation, antibiotic resistance, and social norm emergence. Existing methods largely conflate these two forms and rely on an identical graphical model of selection. We show that this model is valid for static settings but fails to characterize data under evolution, yielding false discovery results. To address this, we introduce a new model that specifically characterizes evolutionary selection, and develop a sound and complete procedure for identifying such models from data across one or multiple environments or generations. Experimental results validate the method's ability to uncover the relevant mechanisms underlying evolution from data.

因果推断进化模型选择机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。