让语言直接控制3D可动物体,支持模糊指令和重复部件操作。
ArtLang: Structured Language-to-Kinematics Grounding for Articulated 3D Actuation

- 用语义-运动图表示物体,结合语言特征实现开放词汇控制。
- 在合成、网格和真实数据上均实现90%以上指令准确率。
- 适合需要自然语言操控3D动画或机器人的人,如设计师与工程师。
可动物体重建虽能恢复显式几何与运动学结构,但其部件常缺乏语义标识,需通过部件索引和数值关节参数控制。我们提出ArtLang,一种对持久重建的可动资产实现开放词汇语言控制的框架。ArtLang将资产表示为语义-运动连接图,并在其表面增强语言特征与图约束运动。开放词汇指令可绑定至已重建部件,允许不确定部件保持未命名状态。类型化解析器将指令转化为包含指代表达、动作、幅度、参考系和关系的指令图。随后求解全局图到图的对齐问题,联合推理语义、空间、关系与运动学兼容性,支持空赋值与模糊情况下的放弃。被接受的指令转换为观察范围内连续关节目标,并通过正向运动学执行。在合成重建、基于网格的资产及真实捕获数据上的实验表明,该方法在重复部件、空间引用、关系命令和模糊指令下均能实现可靠的语义接地与连续可动控制。
原文摘要 · Abstract (English)
Articulated-object reconstructions recover explicit geometry and kinematics, but their parts often remain semantically anonymous and must be controlled through part indices and numerical joint parameters. We present ArtLang, a framework for open-vocabulary language control of persistent reconstructed articulated assets. ArtLang represents an asset as a semantic-kinematic articulation graph and augments its surface with language features and graph-constrained motion. Open-vocabulary proposals are bound to reconstructed parts while allowing uncertain parts to remain unnamed. A typed parser converts a command into a directive graph containing referring expressions, actions, magnitudes, reference frames, and relations. We then solve a global graph-to-graph grounding problem that jointly reasons about semantic, spatial, relational, and kinematic compatibility, with support for null assignments and abstention under ambiguity. Accepted directives are converted into continuous joint targets within the observed motion range and executed through forward kinematics. Experiments on synthetic reconstructions, mesh-based assets, and real captures demonstrate reliable language grounding and continuous articulated control across repeated parts, spatial references, relational commands, and ambiguous instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。