用生物物理过程约束概念瓶颈模型,提升科学预测的可解释性与准确性。
Process-Guided Concept Bottleneck Model
- 基于领域因果机制设计中间概念,引导模型学习符合科学规律的推理路径。
- 在植被生物量估计任务中,误差和偏差均低于多个基准模型。
- 适合需要高可信度、可解释性的科学领域AI应用,如遥感与生态建模。
概念瓶颈模型(CBMs)通过引入中间语义概念提升黑箱深度学习的可解释性。然而,标准CBMs常忽略领域特定关系与因果机制,且依赖完整概念标签,限制了其在科学领域中的应用,因科学数据虽过程明确但标注稀疏。为此,本文提出过程引导的概念瓶颈模型(PG-CBM),通过生物物理上合理的中间概念,强制模型学习遵循领域定义的因果机制。以地表生物量密度估测为例,结果表明PG-CBM在多源异构数据训练下,相比多个基准模型显著降低误差与偏差,同时生成可解释的中间输出。除提升准确性外,该模型还增强透明度,可检测虚假学习,并提供科学洞见,推动科学应用中更可信的AI系统发展。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) improve the explainability of black-box Deep Learning (DL) by introducing intermediate semantic concepts. However, standard CBMs often overlook domain-specific relationships and causal mechanisms, and their dependence on complete concept labels limits applicability in scientific domains where supervision is sparse but processes are well defined. To address this, we propose the Process-Guided Concept Bottleneck Model (PG-CBM), an extension of CBMs which constrains learning to follow domain-defined causal mechanisms through biophysically meaningful intermediate concepts. Using above ground biomass density estimation from Earth Observation data as a case study, we show that PG-CBM reduces error and bias compared to multiple benchmarks, whilst leveraging multi-source heterogeneous training data and producing interpretable intermediate outputs. Beyond improved accuracy, PG-CBM enhances transparency, enables detection of spurious learning, and provides scientific insights, representing a step toward more trustworthy AI systems in scientific applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。