从文字描述生成包含姿态、光照和相机的可编辑3D艺术构图。
Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings

- 联合建模人物姿态、光照与相机参数,实现情感化3D构图生成。
- 在11,911对图文数据上训练,文本检索准确率达32.2%。
- 适合影视概念设计、虚拟制片等需要情绪化视觉参考的场景。
艺术家通过协调人体姿态、光照和相机位置来传达叙事与情感,但现有生成方法通常独立建模这些元素。本文提出文本驱动的可编辑3D构图任务,联合生成人体姿态、主导光照与相机配置,基于情感描述进行输出。研究从2,328幅具象绘画中重建SMPL人体模型,估计低频光照,恢复相机参数,并将每幅画面与ArtEmis描述配对,构建了11,911个文本-构图对。训练了一个支持多角色数量的流匹配变换器,可为同一提示生成多种构图备选。在留出的描述上,模型检索准确率R@1达到32.2%,高于基于CLIP的最近邻检索(16.6%),同时基本保持语料库级别的多样性。结果证明了从文本生成可编辑、情绪化3D构图的可行性。
原文摘要 · Abstract (English)
Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。