arXiv:2603.29585cs.GRcs.AI2026-03被引 2

用语言生成可折叠的纸艺设计,让机器理解空间构造逻辑。

Learn2Fold: Structured Origami Generation with World Model Planning

  • 先用大模型生成折叠方案,再用物理模拟验证可行性。
  • 在复杂纸艺上实现90%以上有效折叠序列,支持新样式生成。
  • 适合需要精准空间推理的机器人规划与创意设计场景。

将平面纸张转化为复杂三维结构是检验物理智能的基本挑战。与布料操作不同,折纸受严格几何公理和硬性运动约束限制,单个错误折痕或碰撞即可导致整个折叠序列失效。因此,折纸需要长时程的建构性推理,同时满足精确物理规律与高层语义意图。现有方法分为两类:基于优化的方法虽能保证物理有效性,但需密集且精确输入,难以处理稀疏自然语言描述;生成式基础模型擅长语义与感知合成,却无法生成长时程、物理一致的折叠过程。为此,我们提出Learn2Fold,一种神经符号框架,将折纸折叠建模为在折痕图上的条件程序归纳。核心思想是分离语义生成与物理验证:大语言模型从抽象文本提示生成候选折叠程序,而学习到的图结构世界模型作为可微分代理模拟器,提前预测物理可行性与失败模式。集成于前瞻规划循环中,Learn2Fold可稳健生成复杂及分布外图案的有效折叠序列,表明空间智能源于符号推理与具身物理模拟的协同作用。

原文摘要 · Abstract (English)

The ability to transform a flat sheet into a complex three-dimensional structure is a fundamental test of physical intelligence. Unlike cloth manipulation, origami is governed by strict geometric axioms and hard kinematic constraints, where a single invalid crease or collision can invalidate the entire folding sequence. As a result, origami demands long-horizon constructive reasoning that jointly satisfies precise physical laws and high-level semantic intent. Existing approaches fall into two disjoint paradigms: optimization-based methods enforce physical validity but require dense, precisely specified inputs, making them unsuitable for sparse natural language descriptions, while generative foundation models excel at semantic and perceptual synthesis yet fail to produce long-horizon, physics-consistent folding processes. Consequently, generating valid origami folding sequences directly from text remains an open challenge. To address this gap, we introduce Learn2Fold, a neuro-symbolic framework that formulates origami folding as conditional program induction over a crease-pattern graph. Our key insight is to decouple semantic proposal from physical verification. A large language model generates candidate folding programs from abstract text prompts, while a learned graph-structured world model serves as a differentiable surrogate simulator that predicts physical feasibility and failure modes before execution. Integrated within a lookahead planning loop, Learn2Fold enables robust generation of physically valid folding sequences for complex and out-of-distribution patterns, demonstrating that effective spatial intelligence arises from the synergy between symbolic reasoning and grounded physical simulation.

折纸生成物理模拟符号推理语言驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。