arXiv:2604.03417cs.LG2026-04

用AI模拟人类审美,让可视化布局更美观高效。

Beauty in the Eye of AI: Aligning LLMs and Vision Models with Human Aesthetics in Network Visualization

  • 用大模型和视觉模型代替人工打标,生成人类偏好的图布局。
  • 优化提示工程后,大模型与人类判断一致率达人与人之间水平。
  • 适合需要大规模美观图布局的科研与工业场景。

网络可视化长期依赖应力等启发式指标,但无单一指标能稳定生成最优效果。数据驱动方法可通过人工偏好标签训练生成模型以逼近人类审美。然而,人工标注成本高、耗时长,现有方法仅基于机器标注数据测试。本文通过27名参与者参与的精心设计用户研究,构建了大规模人类偏好数据集。该数据既用于理解人类审美,也用于训练大语言模型(LLMs)和视觉模型(VMs)作为标注代理。结果表明,结合少样本示例与多格式输入(如图像嵌入)的提示工程显著提升大模型与人类的一致性;进一步通过置信度过滤,使一致性达到人与人之间的水平。此外,经过精心训练的视觉模型也能实现与人类相当的对齐效果。研究证明AI可作为可扩展的人类标注代理。

原文摘要 · Abstract (English)

Network visualization has traditionally relied on heuristic metrics, such as stress, under the assumption that optimizing them leads to aesthetic and informative layouts. However, no single metric consistently produces the most effective results. A data-driven alternative is to learn from human preferences, where annotators select their favored visualization among multiple layouts of the same graphs. These human-preference labels can then be used to train a generative model that approximates human aesthetic preferences. However, obtaining human labels at scale is costly and time-consuming. As a result, this generative approach has so far been tested only with machine-labeled data. In this paper, we explore the use of large language models (LLMs) and vision models (VMs) as proxies for human judgment. Through a carefully designed user study involving 27 participants, we curated a large set of human preference labels. We used this data both to better understand human preferences and to bootstrap LLM/VM labelers. We show that prompt engineering that combines few-shot examples and diverse input formats, such as image embeddings, significantly improves LLM-human alignment, and additional filtering by the confidence score of the LLM pushes the alignment to human-human levels. Furthermore, we demonstrate that carefully trained VMs can achieve VM-human alignment at a level comparable to that between human annotators. Our results suggest that AI can feasibly serve as a scalable proxy for human labelers.

网络可视化大模型人类审美生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。