arXiv:2608.02877cs.LG2026-08

提出新因果发现方法GoT-CD,解决公平性审计因图结构错误而失效的问题。

GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits

  • 用思维链并行生成候选边集,通过确定性验证与硬并集约束构建合法有向无环图
  • 在阿尔茨海默病数据上,8个发现图中有5个丢失敏感属性到结果的不公路径
  • 提醒研究者需结合结构发现进行路径特异性公平分析,避免误判

因果发现从观测数据中恢复有向结构,日益用于临床支持机制推理与预测模型的公平性审计。路径特异性反事实公平性关注受保护属性是否通过非法路径影响结果,但这类评估依赖于给定的因果图,因而继承发现步骤引入的误差。现有发现方法通常以整体结构指标评分,且未评估审计所依赖的具体路径是否在发现中保留,或当该路径缺失时审计报告为何。本文展示,全图思维链推理虽能生成结构上媲美大语言模型基线的无环图,但结构保真度不足以保证公平性审计的可靠性。为此提出GoT-CD:以完整候选边集为推理单元,多图并行生成,经确定性有效性函数评分,通过硬并集约束禁止虚构边,并在提交前用贪心投影确保无环。GoT-CD在五个基准上均返回有效有向无环图,在亚洲、阿尔茨海默病和新冠呼吸系统数据集上优于所有大语言模型方法的有向无环图F1得分。在已知存在不公平路径的阿尔茨海默病基准上,后验路径特异性审计显示,8个发现图中有5个未恢复敏感属性至结果的路径,因此报告整体效应为零,尽管中介效应仍存在,表明必须将路径特异性公平分析与结构发现同步进行。

原文摘要 · Abstract (English)

Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks whether a protected attribute influences an outcome through illegitimate pathways, but these estimands are defined relative to a supplied causal graph and therefore inherit whatever errors the discovery step introduces. Discovery methods are routinely scored on aggregate structural metrics that weight all edges equally, and no established evaluation asks whether the specific pathway an audit depends on survives discovery---or what the audit reports when that pathway is missing. Here we show that full-graph Graph-of-Thoughts reasoning yields acyclic discovered graphs that are structurally competitive with large language model (LLM) baselines, yet that structural fidelity alone does not guarantee fairness-faithful audits. We introduce GoT-CD, in which the reasoning unit is a complete candidate edge set: multiple graphs are generated in parallel, scored by a deterministic validity function, and merged under a hard union constraint that forbids invented edges, with greedy projection enforcing a DAG before commitment. GoT-CD returns a valid DAG on all five reported benchmarks and achieves the best DAG-valid F1 score among LLM methods on Asia, Alzheimer's, and COVID-Respiratory datasets. On an Alzheimer's benchmark with known unfair path, a post-hoc path-specific audit shows that five of eight discovered graphs recover no path from the sensitive attribute to the outcome and therefore report a null overall effect while mediated effects persist, necessitating downstream path-specific fairness analysis along with structural discovery.

因果发现公平性审计图结构路径特异性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。