arXiv:2507.00224cs.CVcs.HC2025-07中稿 · AIED 2025 Late Bre…被引 1

为小组协作中的物体交互,构建了首个6D姿态数据集并验证了现有方法的不足。

Computer Vision for Objects used in Group Work: Challenges and Opportunities

  • 构建了多人协作时小物体6D姿态估计数据集FiboSB
  • 现有算法在小物体与远距离拍摄下检测率不足,平均精度仅0.898
  • 微调YOLO11-x后显著提升检测性能,适合教育场景研究者使用

互动与空间感知技术正在重塑教育框架,尤其在中小学阶段,动手探索有助于深化概念理解。然而,在小组协作任务中,现有系统难以准确捕捉学生与实物之间的交互。这一问题可通过从RGB图像或视频中自动估计物体6D姿态(位置与朝向)来解决。为此,本文提出FiboSB,一个新颖且具有挑战性的6D姿态视频数据集,包含三名参与者协作完成手持小立方体与称重仪任务的场景。该设置因远距离拍摄及物体尺寸小,使6D姿态估计极具难度。我们评估了四种主流6D姿态估计算法在该数据集上的表现,揭示其物体检测模块失效。通过微调YOLO11-x模型,实现mAP_50达0.898。该数据集、基准结果及错误分析为复杂协作场景下的6D姿态估计奠定了基础。

原文摘要 · Abstract (English)

Interactive and spatially aware technologies are transforming educational frameworks, particularly in K-12 settings where hands-on exploration fosters deeper conceptual understanding. However, during collaborative tasks, existing systems often lack the ability to accurately capture real-world interactions between students and physical objects. This issue could be addressed with automatic 6D pose estimation, i.e., estimation of an object's position and orientation in 3D space from RGB images or videos. For collaborative groups that interact with physical objects, 6D pose estimates allow AI systems to relate objects and entities. As part of this work, we introduce FiboSB, a novel and challenging 6D pose video dataset featuring groups of three participants solving an interactive task featuring small hand-held cubes and a weight scale. This setup poses unique challenges for 6D pose because groups are holistically recorded from a distance in order to capture all participants -- this, coupled with the small size of the cubes, makes 6D pose estimation inherently non-trivial. We evaluated four state-of-the-art 6D pose estimation methods on FiboSB, exposing the limitations of current algorithms on collaborative group work. An error analysis of these methods reveals that the 6D pose methods' object detection modules fail. We address this by fine-tuning YOLO11-x for FiboSB, achieving an overall mAP_50 of 0.898. The dataset, benchmark results, and analysis of YOLO11-x errors presented here lay the groundwork for leveraging the estimation of 6D poses in difficult collaborative contexts.

6D姿态教育科技物体检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。