arXiv:2507.08339cs.CL2025-07ACL被引 2

探究金融问答中大模型与推理模型的性能影响因素

What Factors Affect LLMs and RLLMs in Financial Question Answering?

  • 通过模拟长思维链提升大模型表现
  • 推理模型自带长思维链,传统方法效果有限
  • 多语言对齐对推理模型增益小,适合大模型

近期,大语言模型(LLMs)和推理大语言模型(RLLMs)受到广泛关注。RLLMs通过长思维链(Long CoT)过程增强推理能力,显著提升复杂问题解决性能。然而,目前尚无系统性研究探讨如何充分释放LLMs与RLLMs在金融领域的潜力。为此,我们采用五种LLMs和四种RLLMs,评估提示方法、代理框架及多语言对齐方法在金融问答任务中的影响。研究发现:(1) 当前提示方法与代理框架通过模拟长思维链提升了LLMs的金融问答表现;(2) RLLMs具备内在的长思维链能力,导致传统方法对其性能提升有限;(3) 当前先进的多语言对齐方法主要通过延长推理长度改善LLMs的多语言表现,对RLLMs则贡献甚微。此外,我们还讨论了提升两类模型性能的策略,为未来优化提供参考。本研究可为金融领域中LLMs与RLLMs的应用提供重要借鉴。

原文摘要 · Abstract (English)

Recently, large language models (LLMs) and reasoning large language models (RLLMs) have gained considerable attention from many researchers. RLLMs enhance the reasoning capabilities of LLMs through Long Chain-of-Thought (Long CoT) processes, significantly improving the performance of LLMs in addressing complex problems. However, there are few works that systematically explore what methods can fully unlock the performance of LLMs and RLLMs within the financial domain. To investigate the impact of various methods on LLMs and RLLMs, we utilize five LLMs and four RLLMs to assess the effects of prompting methods, agentic frameworks, and multilingual alignment methods on financial question-answering tasks. Our research findings indicate: (1) Current prompting methods and agent frameworks enhance the performance of LLMs in financial question answering by simulating Long CoT; (2) RLLMs possess inherent Long CoT capabilities, which limits the effectiveness of conventional methods in further enhancing their performance; (3) Current advanced multilingual alignment methods primarily improve the multilingual performance of LLMs by extending the reasoning length, which yields minimal benefits for RLLMs. Additionally, we discuss strategies for enhancing the performance of LLMs and RLLMs in financial question answering, which may serve as a inspiration for future improvements. We hope that this study can serve as an important reference for LLMs and RLLMs in the field of financial question answering.

金融问答大模型推理增强多语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。