提出自适应阈值方法与新数据集,提升自动驾驶多任务解释系统的准确性与跨文化适应性。
Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles

- 基于置信度敏感性分析,动态调整多任务决策阈值以优化性能。
- 在新数据集IUST-XAI-AD上,自适应阈值使各类任务F1分数显著提升。
- 适合关注可解释性、跨文化驾驶行为研究的自动驾驶开发者与研究人员。
场景理解是自动驾驶系统的核心,依赖深度学习模型。然而,深度学习模型本质上为黑箱,缺乏透明性与安全性。为提升系统可解释性,多任务视觉理解成为关键,需同时预测多种驾驶行为及其解释,以保障安全并建立人机信任。为构建准确且跨文化的可解释系统,本文引入全面的置信度阈值敏感性分析,评估不同阈值以确定各任务最优决策边界。结果表明,传统固定阈值在多任务场景下表现不佳。通过大量实验验证,所提自适应阈值方法显著提升各任务的F1分数。此外,本文构建了IUST-XAI-AD数据集,包含958张带人类标注驾驶决策与推理的图像,填补了特定驾驶情境评估基准的空白,相比现有数据集更具挑战性。实验显示,置信度阈值敏感性分析能显著提升模型性能,而该数据集揭示了跨文化驾驶行为的重要模式。本工作在方法论与评估工具上均具贡献,推动更可靠、可解释、文化自适应的全球部署自动驾驶系统发展。
原文摘要 · Abstract (English)
Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsically black box models, which lack transparency and safety in autonomous driving. To make these systems transparent, multi-task visual understanding has become crucial for explainable autonomous driving perception systems, where simultaneous prediction of multiple driving behaviors and their underlying explanations is essential for safe navigation and human trust in autonomous vehicles. In order to design an accurate and cross-cultural explainable autonomous driving system, we introduce a comprehensive confidence threshold sensitivity analysis that evaluates various threshold values to identify optimal decision boundaries for different tasks. Our analysis demonstrates that traditional fixed threshold approaches are suboptimal for multi-task scenarios. Through extensive evaluation, we demonstrate that our adaptive threshold selection methodology improves F1-scores across different tasks. In addition, we introduce IUST-XAI-AD, a novel dataset consisting of 958 images with human annotations for driving decisions and corresponding reasoning. This dataset addresses the critical gap in domain-specific evaluation benchmarks for distinct driving contexts and provides a more challenging test environment compared to existing datasets. Experimental results demonstrate that confidence threshold sensitivity analysis can significantly improve model performance, while the introduction of the IUST-XAI-AD dataset reveals important insights about cross-cultural driving behavior patterns. The combined contributions of this work provide both methodological advances and practical evaluation tools that can accelerate the development of more reliable, explainable, and culturally-adaptive autonomous driving systems for global deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。