arXiv:2604.17979cs.IR2026-04中稿 · the 2026 6th Inter…被引 2

在资源受限的中小企业中,模型架构比规模更重要,选对推理方式能显著提升金融问答准确率。

Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints

论文配图:Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
图 1 · 摘自论文原文
  • 对比四种推理架构,发现结构化记忆更适合确定性任务
  • 检索增强在对话式、隐含参考的任务中表现更优
  • 提出动态选择策略,兼顾准确性与部署效率,适合中小企业

人工智能和大语言模型正推动金融分析变革,使自然语言接口成为报告生成、决策支持和自动推理的新方式。然而,针对小中型企业(SME)在成本、精度和合规性约束下的实际工作流,不同基于LLM的推理架构表现仍缺乏实证研究。SME通常面临严重基础设施限制,无云GPU预算、无专职AI团队、无大规模推理能力,因此架构效率成为首要考虑。为此,我们设计了一个明确的SME约束评估场景:所有实验均在本地部署的80亿参数指令微调模型上完成,不依赖云基础设施。系统比较了四种推理架构:基线LLM、检索增强生成(RAG)、结构化长期记忆、记忆增强对话推理,在FinQA与ConvFinQA基准上的结果揭示出一种一致的架构倒置现象:结构化记忆在确定性、操作显式任务中提升精度,而检索方法在对话式、引用隐含场景中表现更优。基于此,我们提出一种混合部署框架,动态选择推理策略,平衡数值准确性、可审计性与基础设施效率,为资源受限环境下的金融AI落地提供可行路径。

原文摘要 · Abstract (English)

The rapid adoption of artificial intelligence (AI) and large language models (LLMs) is transforming financial analytics by enabling natural language interfaces for reporting, decision support, and automated reasoning. However, limited empirical understanding exists regarding how different LLM-based reasoning architectures perform across realistic financial workflows, particularly under the cost, accuracy, and compliance constraints faced by small and medium-sized enterprises (SMEs). SMEs typically operate within severe infrastructure constraints, lacking cloud GPU budgets, dedicated AI teams, and API-scale inference capacity, making architectural efficiency a first-class concern. To ensure practical relevance, we introduce an explicit SME-constrained evaluation setting in which all experiments are conducted using a locally hosted 8B-parameter instruction-tuned model without cloud-scale infrastructure. This design isolates the impact of architectural choices within a realistic deployment environment. We systematically compare four reasoning architectures: baseline LLM, retrieval-augmented generation (RAG), structured long-term memory, and memory-augmented conversational reasoning across both FinQA and ConvFinQA benchmarks. Results reveal a consistent architectural inversion: structured memory improves precision in deterministic, operand-explicit tasks, while retrieval-based approaches outperform memory-centric methods in conversational, reference-implicit settings. Based on these findings, we propose a hybrid deployment framework that dynamically selects reasoning strategies to balance numerical accuracy, auditability, and infrastructure efficiency, providing a practical pathway for financial AI adoption in resource-constrained environments.

金融问答推理架构SME应用高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。