arXiv:2604.19098cs.CLcs.AI2026-04ACL被引 3

首个阿拉伯语金融与合规推理基准,填补中东金融AI空白。

SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning

论文配图:SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
图 1 · 摘自论文原文
  • 构建涵盖7项任务的阿拉伯语金融数据集,覆盖伊斯兰金融合规要求。
  • 20个大模型测试显示,语言能力强的模型在生成任务上表现差,事件因果推理差距最大。
  • 适合研究伊斯兰金融、多语言金融AI及中东数字金融助手的开发者。

英文金融自然语言处理因针对财报分析、市场情绪、表格推理和金融问答的基准而快速发展,但阿拉伯语金融NLP几乎空白,尽管阿拉伯语使用者达4.22亿,海湾主权财富达4.9万亿美元,伊斯兰金融产业规模达4-5万亿美金,需对sukuk、murabaha、takaful等工具进行专门的沙里亚合规审查。我们提出Sahm,首个阿拉伯语金融基准,涵盖7项任务:AAOIFI标准问答、基于教令的问答/选择题、会计与商业考试、金融情绪分析、抽取式摘要和事件-原因推理,包含14,380个来自真实监管、法理和企业来源的专家验证样本。评估20个大模型发现,阿拉伯语流利不等于金融推理能力强:在识别任务中得分达91%的模型,在生成任务中显著下降,事件-原因推理的表现差距最明显(1.89-9.84/10)。我们发布该基准与数据集,以支持可信的阿拉伯语金融助手发展。

原文摘要 · Abstract (English)

English financial NLP has advanced rapidly through benchmarks targeting earnings analysis, market sentiment, tabular reasoning, and financial question answering, yet Arabic financial NLP remains virtually nonexistent, despite 422 million speakers, $4.9 trillion in Gulf sovereign wealth, and a $4-5 trillion Islamic finance industry requiring specialized Shari'ah compliance over instruments like sukuk, murabaha, and takaful. We introduce Sahm, the first Arabic financial benchmark spanning seven tasks: AAOIFI standards QA, fatwa-based QA/MCQ, accounting and business exams, financial sentiment analysis, extractive summarization, and event-cause reasoning, comprising 14,380 expert-verified instances from authentic regulatory, juristic, and corporate sources. Evaluating 20 LLMs, we find Arabic fluency does not imply financial reasoning: models achieving 91% on recognition tasks drop sharply on generation, and event-cause reasoning exposes the widest performance gap (1.89-9.84/10). We release the benchmark and dataset to support trustworthy Arabic financial assistants.

金融AI阿拉伯语伊斯兰金融基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。