arXiv:2512.09847cs.CV2025-12中稿 · WACV 2026

实时识别与预测用户在任务中的挣扎,提升智能辅助系统响应能力。

From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities

  • 将挣扎检测转为在线任务,支持提前预测。
  • 提前2秒预测准确率接近实时检测,性能稳定。
  • 模型可跨任务泛化,适合实际应用部署。

理解人类技能表现对智能辅助系统至关重要,挣扎识别可作为用户困难的自然提示。以往研究集中于离线挣扎分类与定位,而真实场景需要能实时检测并预测挣扎的模型。本文将挣扎定位重构为在线检测任务,并进一步拓展至提前预测,即在挣扎发生前进行预警。采用两种现成模型作为基线,实现在线挣扎检测,帧级mAP达70-80%;提前2秒预测时性能略有下降,但依然保持良好表现。研究还考察了跨任务与活动的泛化能力,以及技能演进的影响。尽管活动层面泛化存在较大领域差异,模型仍比随机基线高出4-20%。特征模型最高运行速度达143 FPS,完整流水线(含特征提取)约20 FPS,满足实时辅助需求。

原文摘要 · Abstract (English)

Understanding human skill performance is essential for intelligent assistive systems, with struggle recognition offering a natural cue for identifying user difficulties. While prior work focuses on offline struggle classification and localization, real-time applications require models capable of detecting and anticipating struggle online. We reformulate struggle localization as an online detection task and further extend it to anticipation, predicting struggle moments before they occur. We adapt two off-the-shelf models as baselines for online struggle detection and anticipation. Online struggle detection achieves 70-80% per-frame mAP, while struggle anticipation up to 2 seconds ahead yields comparable performance with slight drops. We further examine generalization across tasks and activities and analyse the impact of skill evolution. Despite larger domain gaps in activity-level generalization, models still outperform random baselines by 4-20%. Our feature-based models run at up to 143 FPS, and the whole pipeline, including feature extraction, operates at around 20 FPS, sufficient for real-time assistive applications.

实时预测行为理解辅助系统在线检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。