arXiv:2604.19204cs.CYcs.LG2026-04

用聊天记录提升表格预测公平性,降低算法偏见

Auditing LLMs for Algorithmic Fairness in Casenote-Augmented Tabular Prediction

论文配图:Auditing LLMs for Algorithmic Fairness in Casenote-Augmented Tabular Prediction
图 1 · 摘自论文原文
  • 用非营利组织的简短案情笔记增强表格分类
  • 微调后模型准确率提升且偏差降低12%
  • 适合关注社会服务算法公平性的研究者与实践者

大型语言模型在高风险社会服务预测任务中日益受到关注,但其算法公平性尚不明确。本文针对一个真实的住房安置预测任务,审计了基于大模型的表格分类在多类别分类误差上的不公平性。实验发现,通过微调并结合案件笔记摘要,模型在提升准确率的同时,将算法公平性偏差降低了12%。对零样本表格分类引入变量重要性分析的结果显示,公平性改善效果不一。鉴于历史住房分配中的不平等,必须对大模型使用进行审计。尽管案件笔记简短且大量删减,但大模型零样本分类未引入额外文本偏见,仅反映原有表格分类的算法偏见。结合微调与笔记摘要可同时提升准确率与公平性。

原文摘要 · Abstract (English)

LLMs are increasingly being considered for prediction tasks in high-stakes social service settings, but their algorithmic fairness properties in this context are poorly understood. In this short technical report, we audit the algorithmic fairness of LLM-based tabular classification on a real housing placement prediction task, augmented with street outreach casenotes from a nonprofit partner. We audit multi-class classification error disparities. We find that a fine-tuned model augmented with casenote summaries can improve accuracy while reducing algorithmic fairness disparities. We experiment with variable importance improvements to zero-shot tabular classification and find mixed results on resulting algorithmic fairness. Overall, given historical inequities in housing placement, it is crucial to audit LLM use. We find that leveraging LLMs to augment tabular classification with casenote summaries can safely leverage additional text information at low implementation burden. The outreach casenotes are fairly short and heavily redacted. Our assessment is that LLM zero-shot classification does not introduce additional textual biases beyond algorithmic biases in tabular classification. Combining fine-tuning and leveraging casenote summaries can improve accuracy and algorithmic fairness.

大模型审计公平性表格预测社会服务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。