用少量数据提升大模型搭积木的语义与结构准确性
Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning

- 基于模型筛选数据,仅用5%原始数据实现优化
- 新方法在语义、结构、物理三方面超越顶尖模型
- 适合需要高效训练积木推理系统的研究者
基于大语言模型的积木搭建需兼顾语义对齐与物理可行性。本文发现一种数据引发的失效模式——physhack:生成结果虽满足物理约束,却存在几何错位、语义不一致或校准不佳问题。为此提出一种基于模型的数据选择方法,仅使用5%原始数据即可提升多评估者下的语义与结构对齐性。进一步引入PVPO强化学习方法,联合物理合理性与体素空间几何奖励,增强大模型积木生成能力。经测试,使用PVPO训练的DeepSeek-Llama-8B在语义、结构和物理三个维度均优于前沿视觉语言模型Kimi-K3,即使后者拥有真实图像作为参考。
原文摘要 · Abstract (English)
LLM-based LEGO assembly requires both semantic grounding and physical feasibility. In this paper, we identify a data-induced failure mode, physhack, in which generated assemblies satisfy physical-validity constraints while remaining geometrically misaligned, semantically inconsistent, or poorly calibrated. To address this challenge, we propose a model-based data selection approach that uses only 5% of the original training data while improving semantic and structural alignment across multiple independent evaluators. We further introduce PVPO, a reinforcement learning method that couples physical-validity and voxel-space geometric rewards to strengthen the LEGO synthesis capabilities of LLMs. PVPO-trained DeepSeek-Llama-8B outperforms the frontier VLM Kimi-K3 across semantic, structural, and physical evaluation dimensions, even when Kimi-K3 is provided with a ground-truth image as an additional reference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。