arXiv:2604.25949cs.RO2026-04

用手机拍几秒物体,自动生成可部署的感知模型。

FalconApp: Rapid iPhone Deployment of End-to-End Perception via Automatically Labeled Synthetic Data

论文配图:FalconApp: Rapid iPhone Deployment of End-to-End Perception via Automatically Labeled Synthetic Data
图 1 · 摘自论文原文
  • 手机拍摄+自动合成数据,快速构建感知模型
  • 每物体平均20分钟生成数据,30毫秒内完成推理
  • 适合需要快速部署感知能力的机器人开发者

机器人可靠感知依赖大规模标注数据,但真实数据集需大量人工标注且耗时。本文提出FalconApp,一款iPhone应用,通过端到端前后端流程,将手持拍摄的刚性物体视频快速转化为掩码检测与6-DoF位姿估计的感知模块。核心贡献是快速移动端部署管道与逼真自动标注工作流:用户拍摄视频后,FalconApp重建可编辑的GSplat资产,合成多样逼真背景,渲染带真实掩码和位姿的合成图像,训练感知模型,并部署回iPhone前端。五种不同几何与外观的刚性物体实验表明,每物体平均约20分钟生成合成数据并训练完成,设备端端到端延迟约30毫秒,在模拟与真实场景下4/5物体的位姿精度优于PnP基线。

原文摘要 · Abstract (English)

Reliable perception for robotics depends on large-scale labeled data, yet real-world datasets rely on heavy manual annotation and are time-consuming to produce. We present FalconApp, an iPhone app with an end-to-end frontend-backend pipeline that turns a short handheld capture of a rigid object into a perception module for mask detection and 6-DoF pose estimation. Our core contribution is a rapid mobile deployment pipeline paired with a photorealistic auto-labeling workflow: from a user-captured video of an object, FalconApp reconstructs an editable GSplat asset, composites it with diverse photorealistic backgrounds, renders synthetic images with ground-truth masks and poses, trains the perception module, and deploys it back to the iPhone frontend. Experiments across five rigid objects with diverse geometry and appearance show that FalconApp produces usable perception models with about 20 minutes of synthetic-data generation and training per object on average, around 30 ms end-to-end on-device latency on iPhone, and better overall pose accuracy than a PnP baseline on 4 / 5 objects in both simulation and real-world evaluation.

端到端感知移动部署合成数据位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。