arXiv:2410.22118cs.CLcs.AI2024-10NAACL被引 9

推理加速会改变大模型的性别种族偏见,且效果不可预测。

The Impact of Inference Acceleration on Bias of LLMs

论文配图:The Impact of Inference Acceleration on Bias of LLMs
图 1 · 摘自论文原文
  • 测试多种加速技术对模型输出偏见的影响
  • 发现加速后偏见程度显著变化,且无规律可循
  • 提醒需针对具体模型和场景评估加速带来的偏见风险

近年来大型语言模型(LLMs)能力飞速发展,但其推理过程成本高、速度慢。为此,研究者提出了量化、剪枝、缓存等多种加速策略,在降低延迟和成本数倍的同时,保持了主流基准上的预测性能。本文首次系统考察这些加速手段对模型生成内容中人口统计学偏见的影响。通过多维度指标分析,发现加速前后模型输出的偏见水平发生显著变化,且结果复杂且难以预测:同一加速策略在不同模型上可能产生截然不同的偏见影响。该研究强调,必须对经过加速优化的模型进行深入、逐案的偏见评估。

原文摘要 · Abstract (English)

Last few years have seen unprecedented advances in capabilities of Large Language Models (LLMs). These advancements promise to benefit a vast array of application domains. However, due to their immense size, performing inference with LLMs is both costly and slow. Consequently, a plethora of recent work has proposed strategies to enhance inference efficiency, e.g., quantization, pruning, and caching. These acceleration strategies reduce the inference cost and latency, often by several factors, while maintaining much of the predictive performance measured via common benchmarks. In this work, we explore another critical aspect of LLM performance: demographic bias in model generations due to inference acceleration optimizations. Using a wide range of metrics, we probe bias in model outputs from a number of angles. Analysis of outputs before and after inference acceleration shows significant change in bias. Worryingly, these bias effects are complex and unpredictable. A combination of an acceleration strategy and bias type may show little bias change in one model but may lead to a large effect in another. Our results highlight a need for in-depth and case-by-case evaluation of model bias after it has been modified to accelerate inference.

大模型偏见评估推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。