arXiv:2607.26952cs.CL2026-07中稿 · COLM

首个基于真实信用卡协议的金融推理测试,揭示大模型在财务规则理解上的短板。

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

论文配图:Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
图 1 · 摘自论文原文
  • 构建真实信用卡协议的1800个问答数据集,含消费者自然提问形式。
  • 程序思维提示(PoT)显著提升模型表现,尤其改善弱推理模型和闭源系统性能。
  • 错误主因是规则误用而非算术,边缘场景更易出错,影响低收入群体。

我们提出CreditCardQA,首个基于真实信用卡协议的金融素养数值推理基准。该数据集包含1800个问题,涵盖第一人称表达,反映消费者对费用、利息和还款的真实疑问。我们在多种大语言模型和推理模型上评估了链式思维(CoT)与程序思维(PoT)提示策略。总体而言,PoT带来稳定性能提升,尤其对基础推理能力较弱的模型及闭源系统效果显著,缩小了开源与闭源系统间的差距。错误分析显示,失败主要源于金融规则误用、遗漏条件及合同条款误解,而非算术错误。进一步分析表明,对比、条件逻辑和金额约束类问题尤为困难。此外,错误常出现在晚付款罚金、小额余额等边缘情形,这些更可能影响低收入或财务脆弱人群。

原文摘要 · Abstract (English)

We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting. Overall, PoT yields consistent performance gains, particularly for models with weaker baseline reasoning, and narrows gaps between open- and closed-source systems. Through error analysis, we show that failures arise less from arithmetic and more from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. We further analyze question difficulty and find that comparisons, conditional logic, and monetary constraints are especially challenging. We also find that errors often arise in edge cases such as late-payment penalties or small-balance scenarios that are more likely to affect lower-income or financially vulnerable individuals.

金融推理大模型评估信用卡可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。