arXiv:2503.17569cs.LGcs.AI2025-03被引 7

用大模型加速因果发现,提升公平性分析效率

Fairness-Driven LLM-based Causal Discovery with Active Learning and Dynamic Scoring

  • 基于大模型和主动学习,用广度优先策略减少查询量
  • 查询次数从平方级降至线性,显著提升可扩展性
  • 适合关注公平性与因果推断的机器学习研究者

因果发现(CD)在多个科学领域中至关重要,有助于揭示现象背后的因果关系。尽管因果算法在缓解机器学习偏见方面取得进展,但其应用受限于大规模数据带来的高计算成本与复杂性。本文提出一种基于大语言模型(LLM)的因果发现框架,借鉴人类专家的推理方式,采用元数据驱动的方法。通过将成对查询转换为更可扩展的广度优先搜索(BFS)策略,查询数量从变量数的平方级降至线性级,有效解决以往方法的可扩展性问题。该方法结合主动学习(AL)与动态评分机制,根据信息增益优先选择查询,综合互信息、偏相关性和LLM置信度评分,更高效准确地构建因果图。通过对推断出的因果图进行公平性分析,识别敏感属性对结果的直接与间接影响。与基线方法相比,本方法在因果图构建准确性上的优势凸显了精确建模对理解偏见与保障公平性的关键作用。

原文摘要 · Abstract (English)

Causal discovery (CD) plays a pivotal role in numerous scientific fields by clarifying the causal relationships that underlie phenomena observed in diverse disciplines. Despite significant advancements in CD algorithms that enhance bias and fairness analyses in machine learning, their application faces challenges due to the high computational demands and complexities of large-scale data. This paper introduces a framework that leverages Large Language Models (LLMs) for CD, utilizing a metadata-based approach akin to the reasoning processes of human experts. By shifting from pairwise queries to a more scalable breadth-first search (BFS) strategy, the number of required queries is reduced from quadratic to linear in terms of variable count, thereby addressing scalability concerns inherent in previous approaches. This method utilizes an Active Learning (AL) and a Dynamic Scoring Mechanism that prioritizes queries based on their potential information gain, combining mutual information, partial correlation, and LLM confidence scores to refine the causal graph more efficiently and accurately. This BFS query strategy reduces the required number of queries significantly, thereby addressing scalability concerns inherent in previous approaches. This study provides a more scalable and efficient solution for leveraging LLMs in fairness-driven CD, highlighting the effects of the different parameters on performance. We perform fairness analyses on the inferred causal graphs, identifying direct and indirect effects of sensitive attributes on outcomes. A comparison of these analyses against those from graphs produced by baseline methods highlights the importance of accurate causal graph construction in understanding bias and ensuring fairness in machine learning systems.

因果发现大模型公平性主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。