arXiv:2503.00086cs.CVcs.AI2025-03中稿 · TVCG被引 6

CNN在图表关系推理中泛化能力差,仅在视觉编码一致时优于人类。

Generalization of CNNs on Relational Reasoning with Bar Charts

  • 通过扰动条形图视觉编码,测试CNN与人类对长度比的推理能力。
  • 当训练与测试编码不一致时,CNN性能反而低于人类。
  • CNN对无关视觉属性敏感,适合关注模型鲁棒性研究者阅读。

本文系统研究了卷积神经网络(CNN)与人类在条形图关系推理任务中的泛化能力。首先回顾并更新了以往图形感知实验的基准表现。随后,在经典的关系推理任务——估算条形图中条形长度比——上,通过逐步扰动标准可视化形式来测试CNN的泛化性能。我们还开展用户研究,对比了CNN与人类的表现。结果表明,仅当训练与测试数据具有相同视觉编码时,CNN才优于人类;否则可能表现更差。此外,CNN对各类视觉编码的扰动均敏感,无论其是否与目标条形相关;而人类主要受条形长度影响。研究表明,对可视化进行稳健的关系推理对CNN而言仍具挑战性。提升CNN泛化能力需使其更好识别任务相关的视觉特征。

原文摘要 · Abstract (English)

This paper presents a systematic study of the generalization of convolutional neural networks (CNNs) and humans on relational reasoning tasks with bar charts. We first revisit previous experiments on graphical perception and update the benchmark performance of CNNs. We then test the generalization performance of CNNs on a classic relational reasoning task: estimating bar length ratios in a bar chart, by progressively perturbing the standard visualizations. We further conduct a user study to compare the performance of CNNs and humans. Our results show that CNNs outperform humans only when the training and test data have the same visual encodings. Otherwise, they may perform worse. We also find that CNNs are sensitive to perturbations in various visual encodings, regardless of their relevance to the target bars. Yet, humans are mainly influenced by bar lengths. Our study suggests that robust relational reasoning with visualizations is challenging for CNNs. Improving CNNs' generalization performance may require training them to better recognize task-related visual properties.

CNN关系推理可视化泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。