arXiv:2606.08123cs.CVcs.AI2026-06

提出面向驾驶监控的多维度模型评估框架,突破仅看平均分数的局限。

Beyond aggregate scores: Deployment-aware and non-compensatory benchmarking of vision-based eye-state recognition models for driver monitoring

论文配图:Beyond aggregate scores: Deployment-aware and non-compensatory benchmarking of vision-based eye-state recognition models for driver monitoring
图 1 · 摘自论文原文
  • 构建人中心评估框架,分离多维性能与不可补偿的部署门槛。
  • 仅两个模型满足33毫秒延迟要求,均未通过安全筛查,无模型达标。
  • 不同指标下模型排名差异大,强调部署约束必须严格保留失败信息。

针对安全相关的视觉识别模型选择,传统依赖干净数据下的综合性能评分,但鲁棒性、迁移能力、嵌入式延迟和解释可信度可能引发不同偏好。本文提出人类中心基准框架(HCBF),将多维证据与不可补偿的运行准入条件分离。采用主体独立的MRL Eye协议,对六种轻量级卷积与基于Transformer的眼态识别模型进行评估,涵盖确定性图像退化、零样本迁移、参与者安全的目标域训练,以及在RT-BENE数据集上的留一交叉验证,使用NVIDIA Jetson Nano进行TensorRT FP32推理,并以黑盒RISE方法检验解释可信度。干净环境下MRL宏F1得分在0.9566至0.9794之间,而零样本RT-BENE宏F1仅为0.2066至0.7771。目标域适配效果变化范围-0.0406至0.4807,且在不同架构间重新分配了双向误差。仅有MobileNetV3-Large和ShuffleNetV2满足33.333毫秒双目对延迟要求,但两者均未通过预设的安全相关筛选。其余四模型均不达标,导致合格模型集合为空。归一化删除AUC为0.5826至0.9113,归一化插入覆盖率在17.6%至88.6%之间。模型排序在干净预测、抗干扰能力、迁移性能、部署表现、解释可信度及历史评分敏感性等维度上显著变化。研究揭示相对排名、多维偏好与运行准入是截然不同的决策。部署感知的基准测试应保留方向性失败与不确定性,当强制要求无法满足时,允许无模型被选中。

原文摘要 · Abstract (English)

Model selection for safety-relevant visual recognition is often based on clean aggregate performance, although robustness, transfer, embedded latency, and explanation faithfulness may produce different preferences. This study presents a Human-Centered Benchmarking Framework (HCBF) that separates multidimensional evidence from non-compensatory operational eligibility. Six compact convolutional and transformer-oriented eye-state recognition models were evaluated using a subject-disjoint MRL Eye protocol, deterministic image corruptions, zero-shot transfer and participant-safe target-domain training with out-of-fold evaluation on RT-BENE, TensorRT FP32 inference on an NVIDIA Jetson Nano, and black-box RISE faithfulness. Clean MRL Macro-F1 ranged from 0.9566 to 0.9794, whereas zero-shot RT-BENE Macro-F1 ranged from 0.2066 to 0.7771. Matched target-domain effects varied from -0.0406 to 0.4807 and redistributed the two directional errors differently across architectures. Only MobileNetV3-Large and ShuffleNetV2 met the 33.333-ms binocular-pair latency deadline, while neither passed the predefined safety-related screen. The remaining four models failed both requirements, yielding an empty eligible set. Normalized deletion AUC ranged from 0.5826 to 0.9113, while normalized insertion coverage varied from 17.6% to 88.6%. Model ordering changed across clean prediction, corruption robustness, transfer, deployment, faithfulness, and historical score sensitivity. These findings show that relative ranking, multidimensional preference, and operational eligibility are distinct decisions. Deployment-aware benchmarking should preserve directional failures and uncertainty and should allow no model to be selected when mandatory requirements are unmet.

模型评估驾驶监控部署约束多维基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。