arXiv:2501.19339cs.CVcs.CL2025-01被引 2

将文字、表格等全转为像素,测试视觉模型能否统一理解

PixelWorld: How Far Are We from Perceiving Everything as Pixels?

  • 所有信息转成像素输入,用视觉模型统一处理
  • 语义理解表现接近传统分词方法,但推理任务下降明显
  • 适合想简化多模态预处理的研究者

近期智能体语言模型越来越多地需与包含紧密交织的视觉与文本信息的真实环境交互,常通过原始相机像素而非独立处理的图像和分词文本进行。这一趋势凸显了统一感知范式的需求。为此,我们探索将一切感知为像素(PEAP),并提出PixelWorld基准,将自然语言、表格、数学表达和图表等输入转化为共享像素空间。跨多个基准的实验表明,PEAP在语义理解任务上表现可与基于分词的方法媲美,说明视觉变换器能在无需显式分词的情况下部分捕捉全局文本语义。然而,在数学和代码等推理密集型任务中性能显著下降,尽管链式思维提示能部分弥补因缺失符号结构带来的差距。进一步发现,当视觉与文本信息高度融合时,将一切表示为像素可简化预处理流程,并避免跨模态错位。PixelWorld因此提供了一个系统化且实用的框架,用于评估统一视觉-语言模型,并推动像素级多模态学习的进一步研究。

原文摘要 · Abstract (English)

Recent agentic language models increasingly need to interact with real-world environments that contain tightly intertwined visual and textual information, often through raw camera pixels rather than separately processed images and tokenized text. This shift highlights the need for a unified perception paradigm. To investigate this idea, we explore Perceive Everything as Pixels (PEAP) and introduce PixelWorld, a benchmark that renders natural-language, tabular, mathematical, and diagrammatic inputs into a shared pixel space. Experiments across multiple benchmarks show that PEAP achieves comparable performance to token-based approaches on semantic understanding tasks, suggesting that vision transformers can partially capture global textual semantics without explicit tokenization. In contrast, reasoning-intensive tasks such as mathematics and code show notable performance degradation, although Chain-of-Thought prompting helps mitigate this gap by compensating for missing symbolic structure. We further find that when visual and textual information are closely integrated, representing everything as pixels simplifies preprocessing and avoids cross-modal misalignment. PixelWorld thus provides a systematic and practical framework for evaluating unified vision--language models and facilitates further exploration of pixel-based multimodal learning.

多模态视觉模型统一感知像素空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。