arXiv:2510.09846cs.LGcs.AI2025-10

用大模型分析表格数据因果关系,准确率超91%。

CALM: A Causal Analysis Language Model for Tabular Data in Complex Systems with Local Scores, Conditional Independence Tests, and Relation Attributes

  • 基于Mamba架构融合局部得分、独立性检验与关系属性
  • 在仿真和真实生物数据中准确率超91%,优于现有方法
  • 适合需要高精度因果推断的复杂系统研究者

从观测数据中发现因果关系是生物学等科学领域的基础,但现有方法如基于约束的(如PC、causalMGM)和基于评分的(如NOTEARS)存在因果方向难确定、仅限线性关系、对忠实性假设敏感及搜索效率低等问题。尽管大语言模型具备强大推理能力,但其文本设计与表格数据不匹配。为此,我们提出CALM,一种专为复杂系统表格数据设计的因果分析语言模型。CALM采用Mamba架构,通过整合局部因果得分、条件独立性检验和关系属性,捕捉线性、非线性及条件因果机制。在涵盖线性、混合与非线性模型的合成数据集及10个具有严格验证因果关系的真实生物数据集上训练,确保模型鲁棒性和泛化性。实证评估显示,CALM在模拟研究中准确率超过91%,并在丙型肝炎病毒进展的现实应用中成功识别关键因果因素。该工作标志着将语言模型模式识别能力有效应用于表格数据因果发现的重要进展。

原文摘要 · Abstract (English)

Causal discovery from observational data is fundamental to scientific fields like biology, where controlled experiments are often impractical. However, existing methods, including constraint-based (e.g., PC, causalMGM) and score-based approaches (e.g., NOTEARS), face significant limitations. These include an inability to resolve causal direction, restrictions to linear associations, sensitivity to violations of the faithfulness assumption, and inefficiency in searching vast hypothesis spaces. While large language models (LLMs) offer powerful reasoning capabilities, their application is hindered by a fundamental discrepancy: they are designed for text, while most causal data is tabular. To address these challenges, we introduce CALM, a novel causal analysis language model specifically designed for tabular data in complex systems. CALM leverages a Mamba-based architecture to classify causal patterns from pairwise variable relationships. It integrates a comprehensive suite of evidence, including local causal scores, conditional independence tests, and relational attributes, to capture a wide spectrum of linear, nonlinear, and conditional causal mechanisms. Trained on a diverse corpus of synthetic data (from linear, mixed, and nonlinear models) and 10 real-world biological datasets with rigorously validated causal relationships, our model ensures robustness and generalizability. Empirical evaluation demonstrates that CALM significantly outperforms existing methods in both simulation studies, achieving over 91% accuracy, and in a real-world application identifying causal factors in Hepatitis C virus progression. This work represents a significant step towards accurate and generalizable causal discovery by successfully adapting the pattern recognition capabilities of language models to the intricacies of tabular data.

因果发现大模型表格数据生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。