arXiv:2602.07096q-fin.STcs.AI2026-02ACL被引 9

测试大模型在金融问题隐含信息缺失时的推理能力,发现多数模型会盲目作答。

RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

  • 移除金融题目的关键前提,制造信息不足的场景
  • 通用模型答错率升至68%,专业模型也难识别缺条件
  • 适合关注金融AI可信度与风险控制的研究者

可靠的金融推理不仅需要能回答问题,更需知道何时无法作答。现实中,许多金融问题依赖未明说的隐含假设,导致问题看似可解实则信息不足。我们提出REALFIN,一个双语基准,通过系统性地移除金融类考题中的关键前提,保持语言自然,评估模型在三种形式下的表现:直接作答、识别信息缺失、拒绝不合理选项。结果显示,当关键条件缺失时,模型性能显著下降。通用大模型倾向于过度自信并猜测,而多数金融专用模型也无法明确识别缺失前提。这些结果揭示了当前评估体系的重大缺陷,表明可靠的金融模型必须具备‘不回答’的判断力。

原文摘要 · Abstract (English)

Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice, problems often rely on implicit assumptions that are taken for granted rather than stated explicitly, causing problems to appear solvable while lacking enough information for a definite answer. We introduce REALFIN, a bilingual benchmark that evaluates financial reasoning by systematically removing essential premises from exam-style questions while keeping them linguistically plausible. Based on this, we evaluate models under three formulations that test answering, recognizing missing information, and rejecting unjustified options, and find consistent performance drops when key conditions are absent. General-purpose models tend to over-commit and guess, while most finance-specialized models fail to clearly identify missing premises. These results highlight a critical gap in current evaluations and show that reliable financial models must know when a question should not be answered.

金融AI大模型推理能力可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。