LLMs偏爱主流计量系统,导致小众文化用户需更高算力成本。
On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
- 模型默认使用训练数据中的主流计量体系
- 不同计量系统下准确率差异显著,最高相差15个百分点
- 用思维链推理可提升性能,但增加测试时计算量
不同文化使用不同的计量系统(如货币),但其转换关系明确,人类可自由选择表达方式。大型语言模型(LLMs)面向多元文化用户,应能跨系统准确提供信息。我们基于新构建的数据集,测试了七种开源LLM在三种测量类型上的表现,回答三个关键问题:(RQ1) 模型对各类测量的默认系统是什么?(RQ2) 答案与准确率是否随系统变化?(RQ3) 能否通过推理缓解少数系统面临的挑战?结果表明,模型倾向于采用训练数据中占主导地位的计量系统;且在不同系统间表现不稳定,准确率波动达15个百分点以上。尽管使用思维链(CoT)等推理方法可部分缓解此问题,但导致响应长度显著增加,大幅提高测试时计算开销,进一步加剧了小众文化用户的使用障碍。
原文摘要 · Abstract (English)
Measurement systems (e.g., currencies) differ across cultures, but the conversions between them are well defined so that humans can state facts using any measurement system of their choice. Being available to users from diverse cultural backgrounds, large language models (LLMs) should also be able to provide accurate information irrespective of the measurement system at hand. Using newly compiled datasets we test if this is the case for seven open-source LLMs, addressing three key research questions: (RQ1) What is the default system used by LLMs for each type of measurement? (RQ2) Do LLMs' answers and their accuracy vary across different measurement systems? (RQ3) Can LLMs mitigate potential challenges w.r.t. underrepresented systems via reasoning? Our findings show that LLMs default to the measurement system predominantly used in the data. Additionally, we observe considerable instability and variance in performance across different measurement systems. While this instability can in part be mitigated by employing reasoning methods such as chain-of-thought (CoT), this implies longer responses and thereby significantly increases test-time compute (and inference costs), marginalizing users from cultural backgrounds that use underrepresented measurement systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。