AutoDesign让AI自动优化设计流程,实现长期自主改进。
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

- 用元优化器指导代码代理,递归改进设计框架
- 在海报生成任务中达78.32分,优于商用系统7.45分
- 全自动化运行40分钟内完成高质量海报生成,适合研究者使用
将多模态信息转化为结构化媒体输出可视为一个以模型约束系统为核心的长时程智能体过程。理想的约束系统应符合人类设计先验,并通过实证探索积累可复用经验,实现递归自我优化,但现有范式仍为静态模式。本文提出AutoDesign框架,其符合人类设计先验,由元约束优化器引导代码代理基于回放反馈递归改进约束。为实例化与评估该框架,聚焦学术论文转海报任务,构建PosterBench(含100篇跨五学科主赛道论文)及PosterBench-mini(10篇共享子集)。在PosterBench主赛道上,AutoDesign得分为78.32,超越闭源商业系统Claude Design的70.87;在七种代码代理-模型配置中,引入学习到的DesignHarness使平均得分从54.99提升至67.39(+12.4%)。在完全自主的长时程循环中,40分钟内执行253次工具调用与11次编辑,耗资不足3美元,达成人类评估中的会议海报平均质量。盲评实验显示,AutoDesign获最高人类偏好。
原文摘要 · Abstract (English)
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback. To instantiate and evaluate this framework, we focus on the academic paper-to-poster generation task and introduce PosterBench, comprising a 100-paper Main Track spanning five disciplines and PosterBench-mini, a shared 10-paper subset for controlled evaluation. On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points. Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under $3, reaching average conference-poster quality in human evaluation. A system-blind human study further demonstrates that AutoDesign achieves the highest human preference among evaluated systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。