arXiv:2606.19416cs.LG2026-06

评测房贷代理模型表现,发现大模型存在偏见并提出校准方法提升准确率。

MortarBench: Evaluating Mortgage Loan Origination Agents

论文配图:MortarBench: Evaluating Mortgage Loan Origination Agents
图 1 · 摘自论文原文
  • 构建真实数据分布的房贷流程评测基准MortarBench。
  • 顶尖大模型最高仅77.1%准确率,且对非英语姓名存在系统性偏见。
  • 提出CRIT框架,提升至80.5%准确率并改善风险判断与公平性。

贷款发起是贷款机构从申请、审核到批准和放款的关键流程,用于评估申请人资质与风险水平。近年来,企业开始使用房贷代理模型辅助人工贷款专员,但缺乏公开评测基准。为此,我们提出MortarBench——一个房贷代理模型评测基准。该基准通过金融数据合成与变异管道生成覆盖广泛边缘案例的数据,符合真实世界分布与问题特征。实验发现,当前最先进的大语言模型(LLMs)表现不佳,闭源模型最高仅达77.1%的精确匹配准确率。我们还发现,模型对非英语姓名存在系统性偏见。针对此问题,我们提出CRIT——一种置信度校准框架,可将准确率提升至80.5%,同时增强风险控制能力并降低偏见。

原文摘要 · Abstract (English)

Loan origination is the process by which a lender creates a new loan, from application and underwriting through approval and funding. This process serves a critical role in evaluating the eligibility and level of risk posed by an applicant. Recently, firms have begun using mortgage loan agents to augment human loan officers, despite a lack of any public benchmark. To fill this gap, we present MortarBench, a loan origination agent benchmark. MortarBench uses a financial data synthesis and mutation pipeline to generate examples with broad edge case coverage that match real-world distributions and questions. We find that state-of-the-art large language models (LLMs) perform poorly, with closed-source models achieving at most 77.1\% exact match accuracy. We also discover systematic biases in LLM perception of foreignness related to non-English names. Noting these weaknesses, we introduce CRIT, a confidence calibration framework. Our method increases accuracy to 80.5\% while improving risk management steering and reducing bias.

贷款代理大模型评测偏见检测风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。