发现图表转表格中纵轴信息偏差,影响多模态模型表现
Assessing Y-Axis Influence: Bias in Multimodal Language Models on Chart-to-Table Translation

- 构建新框架FairChart2Table分析纵轴相关偏见
- 纵轴刻度长度、数量、范围等显著影响模型准确率
- 在特定模型中提供纵轴信息可大幅提升翻译性能
图表转表格任务将图表图像转化为结构化表格数据,对多模态语言模型(MLM)回答复杂问题至关重要。我们发现公共图表数据集中不同纵轴信息维度的图像数量存在不平衡,这种不平衡可能引入无意偏差,导致MLM性能不均。此前工作未系统考察此类偏差。为此,我们提出新框架FairChart2Table,针对五种前沿模型分析纵轴相关偏见。关键发现:(1) 纵轴刻度值数字长度、主刻度数量、数值范围及刻度格式(如缩写或科学计数法)均存在显著偏差;(2) 图表中图例/实体数量影响MLM表现;(3) 在提示中加入纵轴信息可显著提升部分MLM的性能。
原文摘要 · Abstract (English)
Chart-to-table translation converts chart images into structured tabular data. Accurate translation is crucial for Multimodal Language Model (MLM) to answer complex queries. We observe imbalances in the number of images across different aspects of the y-axis information in public chart datasets. Such imbalances can introduce unintended biases, causing uneven MLM performance. Previous works have not systematically examined these biases. To address this gap, we propose a new framework, FairChart2Table, for analyzing y-axis-related bias on five state-of-the-art models. Key Findings: (1) There are significant y-axis biases related to the digit length of the major tick values, the number of major ticks, the range of values, and the tick value format (e.g., abbreviation or scientific format). (2) The number of legends/entities in chart images impacts MLM performance. (3) Prompting MLM with y-axis information can significantly enhance the performance for some MLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。