提出可解释的三模式人机协作框架,让系统自动决定何时自主处理、转交人类或协同决策。
DeCoDe: Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models
- 基于概念瓶颈模型,用可解释的概念表示做决策
- 在真实数据集上优于纯AI、纯人工及传统延迟方案
- 支持灵活协作模式,适合高风险需透明决策场景
在人机协作中,核心挑战是如何判断任务应由AI独立完成、转交人类专家,还是通过协同互补完成。现有学习延迟方法多为二元选择,忽视了人机互补优势,且缺乏可解释性,难以在高风险场景中让用户理解并修正模型推理。为此,我们提出基于解耦概念瓶颈模型的延时与互补决策框架(DeCoDe),其决策基于人类可理解的概念表示,提升整个过程的透明度。该框架支持三种灵活模式:自主AI预测、转交人类、人机协同互补,由一个以概念级输入为条件的门控网络选择,采用新型代理损失训练,平衡准确率与人力成本。实验表明,DeCoDe在真实数据集上显著优于仅用AI、仅用人及传统延迟基线,在存在噪声标注情况下仍保持强鲁棒性和可解释性。
原文摘要 · Abstract (English)
In human-AI collaboration, a central challenge is deciding whether the AI should handle a task, be deferred to a human expert, or be addressed through collaborative effort. Existing Learning to Defer approaches typically make binary choices between AI and humans, neglecting their complementary strengths. They also lack interpretability, a critical property in high-stakes scenarios where users must understand and, if necessary, correct the model's reasoning. To overcome these limitations, we propose Defer-and-Complement Decision-Making via Decoupled Concept Bottleneck Models (DeCoDe), a concept-driven framework for human-AI collaboration. DeCoDe makes strategy decisions based on human-interpretable concept representations, enhancing transparency throughout the decision process. It supports three flexible modes: autonomous AI prediction, deferral to humans, and human-AI collaborative complementarity, selected via a gating network that takes concept-level inputs and is trained using a novel surrogate loss that balances accuracy and human effort. This approach enables instance-specific, interpretable, and adaptive human-AI collaboration. Experiments on real-world datasets demonstrate that DeCoDe significantly outperforms AI-only, human-only, and traditional deferral baselines, while maintaining strong robustness and interpretability even under noisy expert annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。