用多语言微调提升金融因果关系问答的跨语言性能。
Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026

- 对比三种模型:编码器、编码器-解码器和解码器仅模型,测试多语言适配效果。
- 微调后系统在英语任务中得分4.8140,西班牙语任务中得4.7753,表现领先。
- 适合关注金融文本因果分析与多语言自然语言处理的研究者。
本文介绍团队 HSA_CORAL 参加 FinCausal 2026 共享任务的提交方案,目标是从金融文本中通过抽取式问答提取因果关系,涵盖英语和西班牙语。我们比较了三种建模方法:(i) 使用多语言 BERT 的编码器仅模型进行标记;(ii) 使用多语言 BART 的编码器-解码器生成模型;(iii) 使用提示优化、少量示例和监督微调的解码器仅大模型(Llama 3.1 和 GPT 变体)。实验表明,提示和少量示例已具竞争力,而监督微调带来最大提升。最佳系统为在英文与西班牙语训练数据合并后微调的 GPT-4.1 Mini,其在英语子任务中取得并列最高分(4.8140),在西班牙语任务中排名第三(4.7753),均基于共享任务的 LLM-as-a-judge 评估指标。结果凸显任务特定适配与多语言微调在金融因果问答跨语言迁移中的价值。
原文摘要 · Abstract (English)
This paper describes team HSA_CORAL's submission to the FinCausal 2026 shared task on extracting cause-effect relations from financial narratives via extractive question answering in English and Spanish. We compare three modeling families: (i) encoder-only token tagging with multilingual BERT, (ii) encoder-decoder generation with multilingual BART, and (iii) decoder-only LLMs (Llama 3.1 and GPT variants) using prompt refinement, few-shot demonstrations, and supervised fine-tuning. Across settings, prompting and few-shot examples yield competitive performance, while supervised fine-tuning provides the largest gains. Our best system, GPT-4.1 Mini fine-tuned on combined English and Spanish training data, achieves a tied highest score on the English subtask (score 4.8140) and ranks third on Spanish (score 4.7753) under the shared task's LLM-as-a-judge metric. Overall, the results highlight the value of task-specific adaptation and multilingual fine-tuning for cross-lingual transfer in financial causality QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。