arXiv:2607.22434cs.ROcs.AI2026-07

机器人学会用影子表达手势,实现视觉沟通。

Robot Learning to Communicate through Projected Visual Abstractions

论文配图:Robot Learning to Communicate through Projected Visual Abstractions
图 1 · 摘自论文原文
  • 用可变形软体手建模影子与动作的关系,自动生成目标影子。
  • 通过梯度优化和碰撞检测,实现物理可行的动态影子表演。
  • 适用于手语、皮影戏等视觉表达,适合人机交互研究者。

人类常通过身体的抽象影像(如影子、轮廓、倒影)进行交流,但机器人仍局限于以自身形态表达。本文提出一种具备21自由度灵巧手的机器人系统,通过软性皮肤减少漏光,生成连续的轮廓影像,并利用可微分自模型学习手部姿态与投影影子外观之间的映射关系,基于任务无关的自我探索训练模型。给定目标影子图像或视频时,机器人通过基于梯度的搜索优化手部配置,并在考虑碰撞的仿真中修正,获得物理可行的运动轨迹。针对动态影子表现,引入关键区域目标、时间平滑正则化和关键帧优化,保留视觉重要运动特征的同时降低计算复杂度。实验在仿真与实物平台上验证了该系统在手语手势、手影戏及动物动作模仿中的有效性,为机器人操控自身投影视觉抽象提供了通信与视觉叙事的新框架。

原文摘要 · Abstract (English)

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visually distinct from the body itself. Here, we present a robotic system capable of dynamic shadow expression using a 21-degree-of-freedom dexterous hand with compliant soft skin and a learned shadow self-model. The soft-skinned embodiment reduces light leakage to produce visually continuous silhouettes, while the differentiable self-model learns the mapping between hand configurations and projected shadow appearance through task-agnostic self-exploration. Given a target shadow image or video, the robot optimizes its hand configurations through gradient-based search over 1 the learned self-model and refines the solution through collision-aware simulation to obtain physically feasible motions. For dynamic shadow performance, we further introduce expressive-region objectives, temporal smoothness regularization, and keyframe-based optimization to preserve visually important motion cues while reducing optimization complexity. We demonstrate robotic shadow expression across sign-language gestures, hand-shadow puppetry, and animal motion imitation in both simulation and physical experiments. These results establish a framework for enabling robots to manipulate projected visual abstractions of themselves for communication and visual storytelling.

机器人影子表达人机交互灵巧手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。