arXiv:2603.14505cs.CV2026-03被引 2

让大模型用文字画图,发现生成能力能反哺理解能力。

Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs

  • 用ASCII艺术激活大模型的文本内视觉表达能力
  • 构建7000张高质量图文数据集,支持风格演化
  • 首次实证生成与理解相互促进的循环机制

现有多模态方法将视觉生成视为外部过程,依赖像素渲染或代码执行,忽视了大语言模型(LLMs)内在的视觉表示潜力。本文通过ASCII艺术——一种紧凑、高效且原生文本的视觉格式,解锁这一潜力。提出SVE-ASCII统一框架,用于在纯文本空间中激发并评估符号化视觉表达。为弥补系统性资源匮乏,构建了ASCIIArt-7K数据集,采用创新的“种子与演化”流水线,基于人工标注样本进行上下文风格编辑扩展。进一步实施统一指令微调策略,联合优化文本到ASCII生成与ASCII到文本理解。关键实验揭示任务双重性现象:尽管感知有助于生成,但本研究提供有力证据表明,生成训练显著提升视觉理解能力。这证实了符号化视觉处理中存在相互增强的循环关系,此前虽被推测但极少在视觉领域实证。我们发布数据集、ASCIIArt-Bench基准及SVE-ASCII模型,建立原生文本式视觉智能的坚实基线。

原文摘要 · Abstract (English)

Current multimodal approaches predominantly treat visual generation as an external process, relying on pixel rendering or code execution, thereby overlooking the native visual representation capabilities latent within Large Language Models (LLMs). In this work, we unlock this potential through ASCII art, a compact, efficient, and text-native visual format. We introduce SVE-ASCII, a unified framework designed to elicit and benchmark Symbolic Visual Expression directly within the pure text space. To address the scarcity of systematic resources, we construct ASCIIArt-7K, a high-quality dataset synthesized via a novel "Seed-and-Evolve" pipeline that augments human-curated anchors through in-context stylistic editing. We further implement a unified instruction-tuning strategy that jointly optimizes for both Generation (Text-to-ASCII) and Understanding (ASCII-to-Text). Crucially, our experiments reveal a critical phenomenon regarding task duality: while it is established that perception aids generation, we provide compelling evidence that generative training significantly enhances visual comprehension. This confirms a mutually reinforcing cycle in symbolic visual processing, a relationship previously hypothesized but rarely empirically demonstrated in the visual domain. We release our dataset, the ASCIIArt-Bench benchmark, and the SVE-ASCII model, establishing a robust baseline for native text-based visual intelligence.

视觉表达大模型符号化文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。