arXiv:2601.04589cs.CV2026-01被引 2

让AI理解设计稿分层结构,精准按指令修改图文布局

MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing

  • 用分层推理引擎识别图文层并规划修改
  • 在2万+设计稿上测试,显著优于开源模型
  • 适合需要精细排版编辑的设计师或产品经理

真实世界的设计文档(如海报)具有多层结构,融合装饰、文字与图像。从自然语言指令中编辑此类文档需要细粒度的分层感知与协调能力。现有方法多关注单层图像编辑或多层生成,忽略分层编辑所需的推理机制。为此,我们提出多层文档编辑代理MiLDEAgent,结合强化学习训练的多模态推理器与图像编辑器,实现分层理解与精准修改。为系统评估该任务,我们构建了包含2万余份设计文档与多样化编辑指令的人机协作基准集MiLDEBench,配套任务评估协议MiLDEEval涵盖指令遵循、布局一致性、美学质量与文本渲染四个维度。在14个开源与2个闭源模型上的实验表明,现有方法泛化能力差:开源模型常无法完成任务,闭源模型存在格式错误。相比之下,MiLDEAgent展现出强分层推理与精确编辑能力,显著超越所有开源基线,并达到接近闭源模型的性能,确立了多层文档编辑的首个强基线。

原文摘要 · Abstract (English)

Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grained, layer-aware reasoning to identify relevant layers and coordinate modifications. Prior work largely overlooks multi-layer design document editing, focusing instead on single-layer image editing or multi-layer generation, which assume a flat canvas and lack the reasoning needed to determine what and where to modify. To address this gap, we introduce the Multi-Layer Document Editing Agent (MiLDEAgent), a reasoning-based framework that combines an RL-trained multimodal reasoner for layer-wise understanding with an image editor for targeted modifications. To systematically benchmark this setting, we introduce the MiLDEBench, a human-in-the-loop corpus of over 20K design documents paired with diverse editing instructions. The benchmark is complemented by a task-specific evaluation protocol, MiLDEEval, which spans four dimensions including instruction following, layout consistency, aesthetics, and text rendering. Extensive experiments on 14 open-source and 2 closed-source models reveal that existing approaches fail to generalize: open-source models often cannot complete multi-layer document editing tasks, while closed-source models suffer from format violations. In contrast, MiLDEAgent achieves strong layer-aware reasoning and precise editing, significantly outperforming all open-source baselines and attaining performance comparable to closed-source models, thereby establishing the first strong baseline for multi-layer document editing.

多层编辑设计文档分层推理AI设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。