arXiv:2506.12227cs.LGcs.AI2025-06AAAI被引 3

用大模型辅助发现算法偏见路径,提升公平性审计准确性

Uncovering Bias Paths with LLM-guided Causal Discovery: An Active Learning and Dynamic Scoring Approach

  • 结合大模型语义先验与统计方法,动态优化变量查询顺序
  • 在噪声数据中仍能准确识别性别→教育→收入等关键偏见路径
  • 适合关注算法公平性、需可解释审计的高风险场景研究者

确保机器学习公平性需要理解种族、性别等敏感属性如何因果影响结果。现有因果发现(CD)方法在噪声、混杂或数据污染下难以恢复与公平性相关的路径。大语言模型(LLMs)可通过变量元数据中的语义先验提供互补信号。我们提出一种混合式LLM引导的因果发现框架,将广度优先搜索扩展为带主动学习和动态评分机制。通过融合互信息、偏相关性和LLM置信度的综合得分,优先查询关键变量对,实现更高效、鲁棒的结构发现。为评估公平性敏感性,我们基于UCI Adult数据集构建半合成基准,嵌入领域知情的偏见路径,同时引入噪声和潜在混杂因子。评估显示,在噪声条件下,包括我们提出的主动动态评分变体在内的LLM引导方法,显著优于基线,在恢复全局图结构和关键偏见路径(如sex→education→income)方面表现更优。我们分析了LLM驱动洞察与统计依赖性的互补关系,并讨论其在高风险领域公平性审计中的意义。

原文摘要 · Abstract (English)

Ensuring fairness in machine learning requires understanding how sensitive attributes like race or gender causally influence outcomes. Existing causal discovery (CD) methods often struggle to recover fairness-relevant pathways in the presence of noise, confounding, or data corruption. Large language models (LLMs) offer a complementary signal by leveraging semantic priors from variable metadata. We propose a hybrid LLM-guided CD framework that extends a breadth-first search strategy with active learning and dynamic scoring. Variable pairs are prioritized for querying using a composite score combining mutual information, partial correlation, and LLM confidence, enabling more efficient and robust structure discovery. To evaluate fairness sensitivity, we introduce a semi-synthetic benchmark based on the UCI Adult dataset, embedding domain-informed bias pathways alongside noise and latent confounders. We assess how well CD methods recover both global graph structure and fairness-critical paths (e.g., sex-->education-->income). Our results demonstrate that LLM-guided methods, including our active, dynamically scored variant, outperform baselines in recovering fairness-relevant structure under noisy conditions. We analyze when LLM-driven insights complement statistical dependencies and discuss implications for fairness auditing in high-stakes domains.

因果发现算法公平大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。