用数据域采样法揭示CNN在图表识图中的表现规律
A Rigorous Behavior Assessment of CNNs Using a Data-Domain Sampling Regime
- 设计数据域采样框架,从三方面评估CNN图表理解能力
- 1600万次实验显示:训练测试分布越近,模型越接近人类表现
- 适合关注AI可视化理解能力的开发者与评估研究者
我们提出一种数据域采样方法,用于量化CNN在图表图像中的感知行为。该方法从三个维度评估CNN对柱状图的比例估算能力:对训练-测试分布差异的敏感性、对少量样本的稳定性,以及相对于人类观察者的相对专业度。通过对800个CNN模型进行1600万次试验,以及113名人类参与者6825次试验的分析,得出明确结论:CNN的表现可超越人类,其优劣仅取决于训练与测试数据间的距离。该研究展示了机器在解读可视化图像时的简洁而优雅的行为模式。实验代码与结果可在osf.io/gfqc3获取。
原文摘要 · Abstract (English)
We present a data-domain sampling regime for quantifying CNNs' graphic perception behaviors. This regime lets us evaluate CNNs' ratio estimation ability in bar charts from three perspectives: sensitivity to training-test distribution discrepancies, stability to limited samples, and relative expertise to human observers. After analyzing 16 million trials from 800 CNNs models and 6,825 trials from 113 human participants, we arrived at a simple and actionable conclusion: CNNs can outperform humans and their biases simply depend on the training-test distance. We show evidence of this simple, elegant behavior of the machines when they interpret visualization images. osf.io/gfqc3 provides registration, the code for our sampling regime, and experimental results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。