评测大模型将复杂图像转为可执行代码的能力,发现当前技术仍难保结构完整。
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation
- 构建涵盖1080个样本的综合基准,覆盖多种图像类型与编程语言。
- 顶尖模型在复杂场景中仍难以保持代码结构完整性,错误率高。
- 适合研究多模态生成、代码理解与视觉推理的学者使用。
我们提出Omni-I2C,一个全面的基准,用于评估大型多模态模型(LMMs)将复杂、结构化的数字图形转换为可执行代码的能力。该任务对当前LMMs构成非平凡挑战:需同时具备高保真视觉感知能力以解析复杂的空间层次与符号细节,并拥有精确生成表达能力以合成语法正确且逻辑一致的代码。与传统描述性任务不同,Omni-I2C要求整体理解,任何微小的感知幻觉或编码错误都会导致视觉重建完全失败。该基准包含1080个精心策划的样本,覆盖广泛的主题、图像模态和编程语言。通过引入真实用户来源案例,其涵盖从科学可视化到复杂符号表示的各类数字内容,每项均配有可执行参考代码。为增强评估深度,我们的框架将性能解耦为感知保真度与符号精确度,超越表面准确性,揭示当前LMMs在结构缺陷与推理瓶颈方面的细微问题。评估结果表明,领先模型间存在显著性能差距;即使最先进的模型在复杂场景中也难以维持结构完整性,凸显多模态代码生成仍是重大挑战。数据与代码见https://github.com/MiliLab/Omni-I2C。
原文摘要 · Abstract (English)
We present Omni-I2C, a comprehensive benchmark designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code. We argue that this task represents a non-trivial challenge for the current generation of LMMs: it demands an unprecedented synergy between high-fidelity visual perception -- to parse intricate spatial hierarchies and symbolic details -- and precise generative expression -- to synthesize syntactically sound and logically consistent code. Unlike traditional descriptive tasks, Omni-I2C requires a holistic understanding where any minor perceptual hallucination or coding error leads to a complete failure in visual reconstruction. Omni-I2C features 1080 meticulously curated samples, defined by its breadth across subjects, image modalities, and programming languages. By incorporating authentic user-sourced cases, the benchmark spans a vast spectrum of digital content -- from scientific visualizations to complex symbolic notations -- each paired with executable reference code. To complement this diversity, our evaluation framework provides necessary depth; by decoupling performance into perceptual fidelity and symbolic precision, it transcends surface-level accuracy to expose the granular structural failures and reasoning bottlenecks of current LMMs. Our evaluation reveals a substantial performance gap among leading LMMs; even state-of-the-art models struggle to preserve structural integrity in complex scenarios, underscoring that multimodal code generation remains a formidable challenge. Data and code are available at https://github.com/MiliLab/Omni-I2C.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。