arXiv:2607.19856cs.CLcs.AI2026-07综述被引 1

评测多语言金融问答系统,跨语种准确率最高达97.5%。

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

  • 采用多语言金融选择题测试,覆盖英、中、阿、印四语。
  • 最高准确率达97.5%(英语/阿拉伯语),印度语为92.0%。
  • 顶尖模型普遍使用检索增强与大模型复核机制。

FinMMEval 2026 Task 1 评估英语、中文、阿拉伯语和印地语的多语言金融多项选择题问答能力。任务测试系统在跨语言、跨文字环境下对领域术语、数值理解及金融概念推理的准确判断能力。最终测试集包含800道题,每种语言200题;提交时未提供标准答案,各语言独立按准确率排名。最终排行榜共收录13个英语、11个中文、11个阿拉伯语和10个印地语参赛系统。最高准确率在印地语中为92.0%,英语和阿拉伯语达97.5%,且领先团队在所有语言中均表现优异。所提交系统普遍采用检索增强、直接选项评分、语言特定提示、选择性自一致性、置信度检查及基于大语言模型的审核阶段。

原文摘要 · Abstract (English)

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.

金融问答多语言大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。