arXiv:2607.19867cs.CLcs.AI2026-07综述被引 1

多语言金融问答评测,考察跨语种财务信息理解能力。

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

  • 构建多语言财务文本问答任务,支持英、中、日、西、希五种语言。
  • 测试集含256题,分易难两层,顶尖系统ROUGE-1 F1差距小于1%。
  • 适合关注跨语言金融分析与生成模型的研究者。

FinMMEval 2026 Task 2 评估在多语言证据上进行短答案金融问答的能力。每个测试项包含一个英文问题,以及对应英文、中文、日文、西班牙文和希腊文的财务报表与新闻。参赛系统需以JSONL格式提交每题的简明答案。最终测试集共256个题目,均匀分为易、专家两个难度层级;每层包含四个问题模板,分别在32个公司-报告组上实例化。黄金答案在提交期间保密,系统按宏平均项级ROUGE-1 F1得分排序。最终排行榜收录12个提交结果。最强系统表现极为接近,前四名间ROUGE-1 F1差异不足1个百分点。提交的系统论文涵盖检索增强生成、跨语言证据处理、结构化提示、答案压缩与验证策略。

原文摘要 · Abstract (English)

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

金融问答多语言ROUGE-1评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。