arXiv:2602.01527cs.HCcs.AI2026-02

为机器认知设计可视化,需建立独立于人类的视觉原则。

Toward a Machine Bertin: Why Visualization Needs Design Principles for Machine Cognition

  • 提出人机感知差异的本质是质变而非量变,现有设计无法直接迁移。
  • 机器在图表理解上表现异常:人类易懂的布局对机器无效,反之亦然。
  • 倡导构建面向机器的可视化理论体系,类似人类的Bertin体系。

可视化的设计知识——如有效性排序、编码指南、色彩模型、预注意加工规则——源于六十年来对人类视觉的心理物理学研究。然而,视觉语言模型(VLMs)越来越多地在自动化分析流程中处理图表图像,大量基准测试表明,这些以人类为中心的知识无法直接适用于机器。机器表现出不同的编码性能模式,通过基于块的标记化处理图像,而非整体感知,且对人类无感的设计模式可能失效,而有时在人类难以应对的情况下反而成功。当前主流方法是绕过视觉,将图表转为数据表或结构化文本。本文主张,这种做法忽略了更根本的问题:什么样的视觉表示才真正适合机器认知?本文论证,可视化领域需要将面向机器的视觉设计作为独立研究问题。我们综合了VLM基准、视觉推理研究与可视化素养研究的证据,表明人机感知差异是质性的,非仅量级差异,并批判性审视了主流的回避策略。我们提出区分人类导向与机器导向可视化——并非工程架构,而是一种认识论上的分野——并勾勒出发展该领域实证基础的研究议程,迈出构建‘机器Bertin’的第一步,以补充现有以人类为中心的知识体系。

原文摘要 · Abstract (English)

Visualization's design knowledge-effectiveness rankings, encoding guidelines, color models, preattentive processing rules -- derives from six decades of psychophysical studies of human vision. Yet vision-language models (VLMs) increasingly consume chart images in automated analysis pipelines, and a growing body of benchmark evidence indicates that this human-centered knowledge base does not straightforwardly transfer to machine audiences. Machines exhibit different encoding performance patterns, process images through patch-based tokenization rather than holistic perception, and fail on design patterns that pose no difficulty for humans-while occasionally succeeding where humans struggle. Current approaches address this gap primarily by bypassing vision entirely, converting charts to data tables or structured text. We argue that this response forecloses a more fundamental question: what visual representations would actually serve machine cognition well? This paper makes the case that the visualization field needs to investigate machine-oriented visual design as a distinct research problem. We synthesize evidence from VLM benchmarks, visual reasoning research, and visualization literacy studies to show that the human-machine perceptual divergence is qualitative, not merely quantitative, and critically examine the prevailing bypassing approach. We propose a conceptual distinction between human-oriented and machine-oriented visualization-not as an engineering architecture but as a recognition that different audiences may require fundamentally different design foundations-and outline a research agenda for developing the empirical foundations the field currently lacks: the beginnings of a "machine Bertin" to complement the human-centered knowledge the field already possesses.

可视化机器认知VLM设计原则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。