arXiv:2502.06329cs.CL2025-02被引 8

测试大模型在金融长文本问答中的抗干扰能力,发现顶尖模型仍易幻觉或失效。

Expect the Unexpected: FailSafe Long Context QA for Finance

  • 设计六类金融交互扰动,评估模型对查询和上下文错误的鲁棒性。
  • 最高鲁棒模型在17%测试中失效,最强合规模型幻觉率高达41%。
  • 适合关注金融AI可靠性、需防范幻觉风险的研究者与开发者。

我们提出一个新的长上下文金融基准测试集FailSafeQA,用于评估大语言模型在金融问答系统中面对六种人类交互扰动时的鲁棒性和上下文感知能力。聚焦两类典型场景:查询失败与上下文失败。查询失败中,通过改变查询的领域专业性、完整性和语言准确性进行扰动;上下文失败中,模拟上传低质量、无关或空白文档。采用Qwen2.5-72B-Instruct作为裁判模型,基于细粒度评分标准计算24个现成模型的鲁棒性、上下文一致性与合规性得分。结果显示,尽管部分模型能有效缓解输入扰动,却难以在不产生幻觉的前提下保持可靠输出。值得注意的是,最合规的Palmyra-Fin-128k-Instruct在17%测试用例中预测失效;而最鲁棒的OpenAI o3-mini在41%情况下生成虚构信息。结果表明,即使高性能模型仍有显著提升空间,凸显FailSafeQA作为金融场景可信AI开发工具的价值。数据集已公开于:https://huggingface.co/datasets/Writer/FailSafeQA

原文摘要 · Abstract (English)

We propose a new long-context financial benchmark, FailSafeQA, designed to test the robustness and context-awareness of LLMs against six variations in human-interface interactions in LLM-based query-answer systems within finance. We concentrate on two case studies: Query Failure and Context Failure. In the Query Failure scenario, we perturb the original query to vary in domain expertise, completeness, and linguistic accuracy. In the Context Failure case, we simulate the uploads of degraded, irrelevant, and empty documents. We employ the LLM-as-a-Judge methodology with Qwen2.5-72B-Instruct and use fine-grained rating criteria to define and calculate Robustness, Context Grounding, and Compliance scores for 24 off-the-shelf models. The results suggest that although some models excel at mitigating input perturbations, they must balance robust answering with the ability to refrain from hallucinating. Notably, Palmyra-Fin-128k-Instruct, recognized as the most compliant model, maintained strong baseline performance but encountered challenges in sustaining robust predictions in 17% of test cases. On the other hand, the most robust model, OpenAI o3-mini, fabricated information in 41% of tested cases. The results demonstrate that even high-performing models have significant room for improvement and highlight the role of FailSafeQA as a tool for developing LLMs optimized for dependability in financial applications. The dataset is available at: https://huggingface.co/datasets/Writer/FailSafeQA

金融AI长文本鲁棒性幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。