用专家知识动态融合多种因果发现算法,提升真实场景下的准确性。
Dynamic Expert-Guided Model Averaging for Causal Discovery
- 根据边存在性与方向性差异,分步融合数据驱动结果与专家意见。
- 在干净和噪声数据上均优于强基线,尤其在专家有限时表现更稳。
- 适合缺乏先验知识但有专家可咨询的现实因果推断任务。
因果发现的实际应用面临算法繁多且无明显最优选择的困境。面对众多性能相近的方法,集成策略成为自然选择。然而,真实场景常违背常见因果发现算法的假设,迫使依赖专家知识。受动态请求专家知识及大语言模型作为专家的启发,本文提出一种灵活的模型平均方法,通过选择性地查询专家,集成多种因果发现算法。关键在于区分边的存在性与方向性,从而充分利用数据驱动方法与专家输入的互补优势。同时考虑专家不可靠且访问受限的现实情况,利用算法间的分歧来决定何时调用专家以应对更高不确定性。实验表明,该方法在干净与噪声数据上均持续优于强基线。代码与数据已公开于 https://anonymous.4open.science/r/expert-cd-ensemble-3282/。
原文摘要 · Abstract (English)
Would-be practitioners of causal discovery face a dizzying array of algorithms without a clear best choice. This abundance of competitive methods makes ensembling a natural strategy for practical applications. At the same time, real-world use cases frequently violate the assumptions on which common causal discovery algorithms are based, forcing reliance on expert knowledge. Inspired by recent work on dynamically requested expert knowledge and large language models (LLMs) as experts, we present a flexible model averaging method that integrates selective expert querying to ensemble a diverse set of causal discovery algorithms. Crucially, we distinguish between edge existence and orientation, enabling the method to leverage the complementary strengths of data-driven discovery and expert input. We further consider the realistic setting of limited access to an imperfect expert, using disagreement among algorithms to query the expert in cases of greater uncertainty. Experiments demonstrate that our method consistently outperforms strong baselines on both clean and noisy data. Code and data are available at https://anonymous.4open.science/r/expert-cd-ensemble-3282/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。