用合成数据和新评估方法提升金融问答模型的推理能力
From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

- 构建高质合成数据管道,确保问答对准确相关
- 新评估指标按计算表达式匹配,准确反映推理能力
- 结合语义相似性损失,显著提升小模型在金融数据上的表现
金融问答(QA)已成为评估大语言模型在表格、图表和复杂文本等多格式领域任务中表现的关键基准。尽管近期进展使模型具备跨模态推理与多步运算能力,但性能一致性与评估可靠性仍存不足。传统精确匹配(EM)指标常忽略单位或格式微小差异,导致评估失真。本文提出一个综合性改进方案:通过高质量合成数据生成与量化低秩适配(QLoRA)微调小型语言模型(SLMs)。管道包含严格的合成问答对验证机制,确保生成结果正确且相关。提出一种新评估指标,以模型计算出的算术表达式匹配为目标,而非直接比对真实答案,更真实反映推理能力。同时设计改进损失函数,结合语义相似性、新评估指标与标准交叉熵,有效提升模型性能。在基准数据集ConvFinQA上,使用合成数据和所提损失函数微调后,问答准确率显著提升。
原文摘要 · Abstract (English)
Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent advancements have enabled models to reason across modalities and perform multi-step arithmetic operations, limitations remain in performance consistency, and evaluation reliability. In particular, standard evaluation metrics like Exact Match (EM) often fail to account for minor variations such as differences in units or formats, misleading performance assessments. In this work, we propose a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA). Our pipeline includes aggressive data validation for synthetic question answer generation to ensure the relevance and correctness of synthetic question-answer pairs. We introduce a novel evaluation metric that matches answers computed from arithmetic expressions rather than ground-truth answers; providing a more accurate reflection of model reasoning capability. Furthermore, we propose a modified loss function that aligns predicted and reference expressions using semantic similarity, our novel evaluation metric and standard cross-entropy, resulting in improved performance. Experimental results on benchmark datasets, ConvFinQA demonstrate significant gains in QA accuracy after fine-tuning using synthetic dataset and proposed loss function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。