arXiv:2505.04670cs.SEcs.AI2025-05被引 6

首个评估LLM修改TikZ代码视觉效果的基准,解决自然语言改图难题

LLM Code Customization with Visual Results: A Benchmark on TikZ

  • 构建vTikZ基准,用可视化反馈评估代码修改
  • 现有LLM在视觉意图对齐上表现不佳,准确率不足50%
  • 适合关注AI辅助设计、图像生成与编程的科研人员

随着基于AI的代码生成兴起,仅通过自然语言指令修改已有代码以改变视觉结果(如图表或图像)已成为可能,有望降低对深层编程知识的需求。然而,即使经验丰富的开发者也常面临挑战,因需精准定位代码区域、生成合法代码变体,并确保修改可靠符合用户意图。本文提出vTikZ,首个专门评估大型语言模型(LLMs)在保持视觉连贯性前提下定制代码能力的基准。该基准包含精心设计的vTikZ编辑场景、参数化真实答案及利用视觉反馈的评审工具。对前沿LLMs的实证评估显示,现有方案在视觉意图对齐方面表现不佳,暴露出当前AI辅助代码编辑方法的短板。我们主张,vTikZ为整合LLMs与视觉反馈机制开辟新研究方向,可推广至图像处理、艺术创作、网页设计和3D建模等多个领域。

原文摘要 · Abstract (English)

With the rise of AI-based code generation, customizing existing code out of natural language instructions to modify visual results -such as figures or images -has become possible, promising to reduce the need for deep programming expertise. However, even experienced developers can struggle with this task, as it requires identifying relevant code regions (feature location), generating valid code variants, and ensuring the modifications reliably align with user intent. In this paper, we introduce vTikZ, the first benchmark designed to evaluate the ability of Large Language Models (LLMs) to customize code while preserving coherent visual outcomes. Our benchmark consists of carefully curated vTikZ editing scenarios, parameterized ground truths, and a reviewing tool that leverages visual feedback to assess correctness. Empirical evaluation with stateof-the-art LLMs shows that existing solutions struggle to reliably modify code in alignment with visual intent, highlighting a gap in current AI-assisted code editing approaches. We argue that vTikZ opens new research directions for integrating LLMs with visual feedback mechanisms to improve code customization tasks in various domains beyond TikZ, including image processing, art creation, Web design, and 3D modeling.

代码生成视觉反馈LLMTikZ

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。