用无人机实拍视频构建3D人体姿态数据集,提升远距离识人意图能力
Learning to Understand Body Language from Flight through Robust 3D Avatar Placing

- 基于单目深度流生成局部几何世界模型,实现精准3D角色定位
- 在12种架构上训练,真实、重定向和生成动作的意图识别准确率显著提升
- 适用于无人机视觉理解、人机交互等需要远距离感知意图的场景
远距离感知人类运动与意图是社交智能飞行机器人的关键前提,但相关数据极为稀缺。我们提出 Drones2BodyLanguage,一个基于真实无人机画面的3D人体姿态数据集:将表达十种沟通意图的虚拟角色,以精确的位置、尺度和朝向放置于未经修改的4K无人机场景中,并在数百帧相机运动中保持一致。实现该数据集依赖于一个轻量级的局部场景几何世界模型——通过流式单目深度提取语义锚点并升维至3D,利用仿射锚点组合预测放置点,权重具有刚性不变性;再通过SVD拟合地面旋转进行重渲染。在12种架构的场景与运动不交叠划分下,使用该放置数据训练后,真实、重定向及生成动作的意图识别准确率均大幅提高,且在两个真实野外场景中得到验证。
原文摘要 · Abstract (English)
Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion. Enabling it is a lightweight geometric world model of the local scene - semantically selected anchors lifted to 3D through streaming monocular depth - in which a placement point is predicted as an affine anchor combination with provably rigid-invariant weights, and re-rendered under an SVD-fitted ground rotation. Across twelve architectures on scene- and motion-disjoint splits, training on placed data lifts mean intent accuracy by a wide margin for real, retargeted and generated motion alike, with gains confirmed on two in-the-wild scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。