提出可解释的多模态行为建模框架,让AI决策过程像人一样看得懂。
Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
- 用注意力引导概念模型,自动找出影响判断的关键行为特征
- 在表情识别数据集上保持高准确率,同时提供可理解的解释
- 适合需要透明决策的智能系统,如情绪识别、人机交互场景
在智能互联时代,情感计算(AC)使系统能够识别、理解并响应人类行为状态,已成为众多AI系统的核心组成部分。作为负责任AI与以人为本系统可信性的关键,可解释性成为AC领域的核心关注点。特别是欧盟《通用数据保护条例》要求高风险AI系统具备充分可解释性,包括广泛应用于情感计算领域的生物特征与情绪识别系统。现有可解释方法常在可解释性与性能间权衡,多数仅聚焦于突出网络参数,缺乏对利益相关方有意义的领域特定解释,且难以有效融合多模态数据进行联合学习与解释。为此,我们提出一种新颖通用的框架——注意力引导概念模型(AGCM),通过识别导致预测结果的概念及其观测位置,提供可学习的概念化解释。AGCM可通过多模态概念对齐与联合学习扩展至任意时空信号,使利益相关方获得更深入的模型决策洞察。我们在经典面部表情识别基准数据集上验证了AGCM的有效性,并展示了其在更复杂的现实世界人类行为理解应用中的泛化能力。
原文摘要 · Abstract (English)
In the contemporary era of intelligent connectivity, Affective Computing (AC), which enables systems to recognize, interpret, and respond to human behavior states, has become an integrated part of many AI systems. As one of the most critical components of responsible AI and trustworthiness in all human-centered systems, explainability has been a major concern in AC. Particularly, the recently released EU General Data Protection Regulation requires any high-risk AI systems to be sufficiently interpretable, including biometric-based systems and emotion recognition systems widely used in the affective computing field. Existing explainable methods often compromise between interpretability and performance. Most of them focus only on highlighting key network parameters without offering meaningful, domain-specific explanations to the stakeholders. Additionally, they also face challenges in effectively co-learning and explaining insights from multimodal data sources. To address these limitations, we propose a novel and generalizable framework, namely the Attention-Guided Concept Model (AGCM), which provides learnable conceptual explanations by identifying what concepts that lead to the predictions and where they are observed. AGCM is extendable to any spatial and temporal signals through multimodal concept alignment and co-learning, empowering stakeholders with deeper insights into the model's decision-making process. We validate the efficiency of AGCM on well-established Facial Expression Recognition benchmark datasets while also demonstrating its generalizability on more complex real-world human behavior understanding applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。