arXiv:2508.21512cs.LGcs.CL2025-08EMNLP

不同数据转文本方式影响大模型贷款审批的准确与公平性

Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches

  • 对比三种数据序列化方法对大模型性能的影响
  • 上下文学习使准确率提升4.9%至59.6%,但公平性波动大
  • 提醒需关注数据表示方式与公平性设计,适合金融决策研究者

大型语言模型(LLMs)在贷款审批等高风险决策任务中应用日益广泛。然而,它们处理表格数据能力有限,难以兼顾公平性与可靠性。本文评估了三种地理区域(加纳、德国、美国)贷款审批数据集在不同序列化格式下的模型表现与公平性。序列化指将表格数据转换为适合大模型处理的文本格式。结果显示,序列化方式显著影响模型性能与公平性:如GReat和LIFT格式虽提升F1分数,却加剧公平性偏差。值得注意的是,上下文学习(ICL)相比零样本基线使性能提升4.9%至59.6%,但对公平性的影响在不同数据集间差异显著。本研究强调有效表格数据表示方法与公平性感知模型对提升金融决策中大模型可靠性的关键作用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly employed in high-stakes decision-making tasks, such as loan approvals. While their applications expand across domains, LLMs struggle to process tabular data, ensuring fairness and delivering reliable predictions. In this work, we assess the performance and fairness of LLMs on serialized loan approval datasets from three geographically distinct regions: Ghana, Germany, and the United States. Our evaluation focuses on the model's zero-shot and in-context learning (ICL) capabilities. Our results reveal that the choice of serialization (Serialization refers to the process of converting tabular data into text formats suitable for processing by LLMs.) format significantly affects both performance and fairness in LLMs, with certain formats such as GReat and LIFT yielding higher F1 scores but exacerbating fairness disparities. Notably, while ICL improved model performance by 4.9-59.6% relative to zero-shot baselines, its effect on fairness varied considerably across datasets. Our work underscores the importance of effective tabular data representation methods and fairness-aware models to improve the reliability of LLMs in financial decision-making.

大模型贷款审批公平性序列化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。