arXiv:2606.02789cs.CV2026-06

诊断教育场景中人机交互检测失败原因并针对性优化模型。

Diagnosis of Human Object Interaction Detectors for Real World Educational Applications

  • 构建细粒度错误分类体系,定位真实教学视频中的交互识别问题。
  • 在CCATT数据集上将模型宏平均F1从48.6提升至90.2。
  • 适合需高精度行为分析的虚拟医疗训练等真实教育场景。

人机交互(HOI)识别对于自动分析复杂教育环境中的学生行为至关重要。尽管当前最先进(SOTA)的HOI检测器在基准数据集上表现良好,但在实际教学环境中因特定领域物体、遮挡和复杂视觉条件导致性能下降。本文提出一种诊断驱动框架,结合三元组级HOI错误分类体系与误差因素归因分析,针对混合现实医疗培训中的关键护理空中运输团队(CCATT)场景开展研究。基于对HOI失效模式及其成因的分析,我们设计了基于诊断信息的微调策略,用于将预训练的HOI模型适配至目标域。在CCATT数据集上的实验表明,该方法通过针对诊断出的误差因素进行定向优化,使预训练的CDN模型的宏平均F1从48.6提升至90.2。结果凸显了细致诊断分析在指导真实教育场景中HOI模型精准适应方面的价值。

原文摘要 · Abstract (English)

Human-object interaction (HOI) recognition is critical for automatically analyzing student behavior in complex educational environments. Although state-of-the-art (SOTA) HOI detectors perform well on benchmark datasets, their performance often degrades when deployed in real-world training environments due to domain-specific objects, occlusions, and complex visual conditions. In this paper, we introduce a diagnosis-driven framework that integrates a triplet-level HOI error taxonomy with error-factor attribution analysis for real-world educational video data. We study this problem in the context of Critical Care Air Transport Team (CCATT) mixed-reality medical training. Based on an analysis of HOI failure modes and their causes, we develop a diagnosis-informed refinement strategy for adapting pretrained HOI models to the target domain. Experiments on the CCATT dataset show that this approach improves the macro-F1 score of a pretrained CDN model from 48.6 to 90.2 through targeted refinement guided by diagnosed error factors. These results highlight the value of detailed diagnostic analysis for informing targeted adaptation of HOI models in real-world educational environments.

人机交互行为识别教育应用模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。