针对交警手势识别的可靠性问题,提出动态加权与因果推理结合的新模型。
RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

- 基于姿态置信度动态加权图推理,弱化不可靠关节影响
- 在13.4万帧数据上达到93.33%准确率,优于现有方法3-4个百分点
- 适合自动驾驶中对早期、稳定、鲁棒手势识别有要求的场景
交通警察手势是自动驾驶中的关键感知信号。部署式识别器需从连续视频中因果推断指令,在手臂运动过渡期保持稳定,并避免过度依赖受损的姿态测量。本文提出RSC-GestureNet,一种可靠性感知的选择性因果识别器,将姿态置信度作为首要信号:不可靠关节在图推理中被降权,时间证据以因果方式聚合,通过可靠性感知推理规则选择性输出预测。我们还引入CTPGesture-C,一个可复现的特征级退化基准,包含七类姿态/RGB退化,并设计了基于MediaPipe重处理的RGB级诊断流程。在完整的官方CTPGesture v1数据集(134,424帧标注帧,33,451个因果窗口)上,RSC-GestureNet达到93.33±0.24%准确率,91.71±0.27%宏平均F1,91.69±0.29%在线宏平均F1,98.80±0.07% Early@10,0.153±0.013秒TTC,且在鲁棒宏平均F1上表现最佳。相比复现的交通专用MD-GCN和HLP-GCN基线,其宏平均F1提升3.23-4.11点,在线F1提升2.15-3.07点。结合校准、选择性风险、统计、自适应分支和图像级重提取分析,表明显式建模姿态可靠性能显著提升早期、稳定、鲁棒的交通指令识别能力。
原文摘要 · Abstract (English)
Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain stable around transitional arm motion, and avoid over-trusting corrupted pose measurements. This study presents RSC-GestureNet, a reliability-aware selective causal recognizer, for Chinese traffic police gestures. The model treats pose confidence as a first-class signal: unreliable joints are down weighted during graph reasoning, temporal evidence is aggregated causally, and calibrated predictions are selectively emitted through a reliability-aware inference rule. We further introduce CTPGesture-C, a reproducible feature-level corruption benchmark with seven pose/RGB degradation families, and an RGB-level diagnostic in which corrupted frames are reprocessed by MediaPipe before recognition. On the complete official CTPGesture v1 split (134,424 labeled frames and 33,451 causal windows), RSC-GestureNet achieves 93.33+-0.24% accuracy, 91.71+-0.27% macro-F1, 91.69+-0.29% online macro-F1, 98.80+-0.07% Early@10, 0.153+-0.013 s TTC, and the best robust macro-F1 among evaluated methods. Under the same split and causal protocol, it exceeds reproduced traffic-specific MD-GCN and HLP-GCN baselines by 3.23-4.11 macro-F1 points and 2.15-3.07 online-F1 points. These results, together with calibration, selective-risk, statistical, adaptive-branching, and image-level re-extraction analyses, indicate that explicit pose-reliability modeling improves early, stable, and robust traffic-command recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。