用姿态识别支持规则驱动的多媒体决策,以击剑判罚为例。
FERA: A Pose-Based Framework for Rule-Grounded Multimedia Decision Support with a Foil Fencing Case Study
- 基于姿态追踪与运动建模,将动作转化为可验算的规则输入。
- 在击剑判罚任务中达0.624准确率与0.632宏平均F1。
- 适合需要可解释性、可审计性的实时决策场景。
多媒体决策支持不仅需要识别,更需能被规则检验、人工审计并供下游逻辑使用的显式状态估计。本文提出击剑裁判助手(FERA),一种基于姿态的规则驱动决策框架,并以击剑中的双人快速移动与优先权规则为案例研究。该框架包含:标准参与者追踪、运动特征编码、校准的时间感知、紧凑的结构化决策层及面向解释的检索界面。我们还发布了经审核的基准数据集,包含标注明确且固定划分的训练/测试集,支持可复现评估。在统一协议下,轻量级深度旁路模块提升了最优图模型表现;基于二维特征流的紧凑结构分类器在最终左/右/无判定任务中达到0.624准确率与0.632宏平均F1分数。案例研究揭示了通用设计原则:明确感知与规则应用的边界,保留不确定性,并根据下游需求选择感知前端。
原文摘要 · Abstract (English)
Multimedia decision support requires more than recognition; it requires explicit state estimates that can be checked against rules, audited by humans, and consumed by downstream decision logic. We present the FEncing Referee Assistant (FERA), a pose-based framework for this setting, and study it through foil fencing, where decisions depend on fast bilateral motion and right-of-way rules. The framework separates canonical participant tracking, kinematic tokenization, calibrated temporal perception, a compact structured decision layer, and an explanation-oriented retrieval interface. We also release an audited benchmark with adjudicated labels and fixed folds for reproducible evaluation. Under a shared protocol, a lightweight lifted-depth sidecar strengthens the best graph-based perception model, while a compact structured classifier on the fixed two-dimensional token stream reaches 0.624 accuracy and a 0.632 macro-averaged F1 score on the final Left / Right / None decision. The case study supports a broader design lesson: keep the boundary between perception and rule application explicit, preserve uncertainty, and choose the perception front end according to the downstream operating point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。