用难度自适应课程学习提升2D图像旋转回归的半监督效果
HACMatch Semi-Supervised Rotation Regression with Hardness-Aware Curriculum Pseudo Labeling
- 根据样本难易度动态筛选伪标签,避免固定阈值误判
- 在低标注数据下性能超越现有方法,PASCAL3D+上提升6.2%
- 设计结构化数据增强,保留几何一致性同时增加特征多样性
从2D图像回归3D物体旋转是一项关键且具挑战性的任务,广泛应用于自动驾驶、虚拟现实和机器人控制。现有模型通常依赖大量标注数据或额外信息(如点云、CAD模型)。因此,仅用少量标注2D图像实现半监督旋转回归具有重要价值。虽然近期工作FisherMatch引入了半监督学习,但其基于熵的伪标签过滤机制僵化,难以区分可靠与不可靠样本。为此,我们提出一种难度感知的课程学习框架,动态依据样本难易度选择伪标签,由易到难逐步推进。设计多阶段与自适应课程策略,替代固定阈值过滤,更具灵活性。此外,提出一种专为旋转估计设计的结构化数据增强方法,通过拼接增强后的图像块生成复合图像,在保持关键几何完整性的同时引入特征多样性。在PASCAL3D+和ObjectNet3D上的全面实验表明,该方法在低数据条件下显著优于现有监督与半监督基线,验证了课程学习框架与结构化增强的有效性。
原文摘要 · Abstract (English)
Regressing 3D rotations of objects from 2D images is a crucial yet challenging task, with broad applications in autonomous driving, virtual reality, and robotic control. Existing rotation regression models often rely on large amounts of labeled data for training or require additional information beyond 2D images, such as point clouds or CAD models. Therefore, exploring semi-supervised rotation regression using only a limited number of labeled 2D images is highly valuable. While recent work FisherMatch introduces semi-supervised learning to rotation regression, it suffers from rigid entropy-based pseudo-label filtering that fails to effectively distinguish between reliable and unreliable unlabeled samples. To address this limitation, we propose a hardness-aware curriculum learning framework that dynamically selects pseudo-labeled samples based on their difficulty, progressing from easy to complex examples. We introduce both multi-stage and adaptive curriculum strategies to replace fixed-threshold filtering with more flexible, hardness-aware mechanisms. Additionally, we present a novel structured data augmentation strategy specifically tailored for rotation estimation, which assembles composite images from augmented patches to introduce feature diversity while preserving critical geometric integrity. Comprehensive experiments on PASCAL3D+ and ObjectNet3D demonstrate that our method outperforms existing supervised and semi-supervised baselines, particularly in low-data regimes, validating the effectiveness of our curriculum learning framework and structured augmentation approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。