发现大模型对不同数据格式存在系统性偏见,影响多源信息融合公平性。
Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
- 通过三阶段实证分析揭示格式偏见普遍存在
- 偏见由信息丰富度、结构质量等数据因素驱动
- 偏见源于注意力分配不均,可借重加权干预缓解
大型语言模型(LLMs)在处理文本、表格、信息框和知识图谱等异构数据时日益重要,但其对特定格式的系统性偏见可能破坏多源数据的公平整合,引发推理错误并增加下游任务风险。本文首次通过三阶段实证研究系统分析了格式偏见:第一阶段检测多种模型中的偏见存在与方向;第二阶段探究信息丰富度、结构质量、表示类型等数据层面因素的影响;第三阶段分析偏见在模型注意力模式中的表现,并测试轻量级干预的有效性。结果表明,格式偏见在不同模型家族中具有一致性,主要受数据特征驱动,并与模型内部注意力不平衡密切相关。基于此,提出三个未来方向:通过格式修复与标准化提升数据预处理质量,引入推理时注意力重加权等干预手段,构建格式平衡的训练语料库。这些方向将推动更鲁棒、更公平的异构数据处理系统发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly employed in applications that require processing information from heterogeneous formats, including texts, tables, infoboxes, and knowledge graphs. However, systematic biases toward particular formats may undermine LLMs' ability to integrate heterogeneous data impartially, potentially resulting in reasoning errors and increased risks in downstream tasks. Yet it remains unclear whether such biases are systematic, which data-level factors drive them, and what internal mechanisms underlie their emergence. In this paper, we present the first comprehensive study of format bias in LLMs through a three-stage empirical analysis. The first stage explores the presence and direction of bias across a diverse range of LLMs. The second stage examines how key data-level factors influence these biases. The third stage analyzes how format bias emerges within LLMs' attention patterns and evaluates a lightweight intervention to test its effectiveness. Our results show that format bias is consistent across model families, driven by information richness, structure quality, and representation type, and is closely associated with attention imbalance within the LLMs. Based on these investigations, we identify three future research directions to reduce format bias: enhancing data pre-processing through format repair and normalization, introducing inference-time interventions such as attention re-weighting, and developing format-balanced training corpora. These directions will support the design of more robust and fair heterogeneous data processing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。