arXiv:2509.11365cs.CL2025-09被引 1

通过提示工程与集成学习,提升大模型在阿拉伯语临床问答中的表现。

!MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering through Prompt Engineering and Ensemble Learning

  • 采用少样本提示+三组提示集成,优化分类准确率。
  • 在多项任务中取得第二名,涵盖选择题与开放问答。
  • 适合关注多语言医疗AI与提示工程的开发者。

我们提交了 AraHealthQA-2025 共享任务第2赛道(通用阿拉伯语健康问答,MedArabiQ)的系统,方法在子任务1(选择题问答)和子任务2(开放问答)中均获得第二名。针对子任务1,我们使用 Gemini 2.5 Flash 模型,结合少样本提示、数据预处理及三组提示配置的集成,提升标准题、偏见题与填空题的分类准确率。针对子任务2,采用统一提示策略,包含角色扮演(阿拉伯语医学专家)、少样本示例与后处理,生成简洁响应,覆盖填空、医患问答、语法纠错与改写变体等场景。

原文摘要 · Abstract (English)

We present our systems for Track 2 (General Arabic Health QA, MedArabiQ) of the AraHealthQA-2025 shared task, where our methodology secured 2nd place in both Sub-Task 1 (multiple-choice question answering) and Sub-Task 2 (open-ended question answering) in Arabic clinical contexts. For Sub-Task 1, we leverage the Gemini 2.5 Flash model with few-shot prompting, dataset preprocessing, and an ensemble of three prompt configurations to improve classification accuracy on standard, biased, and fill-in-the-blank questions. For Sub-Task 2, we employ a unified prompt with the same model, incorporating role-playing as an Arabic medical expert, few-shot examples, and post-processing to generate concise responses across fill-in-the-blank, patient-doctor Q&A, GEC, and paraphrased variants.

医疗问答阿拉伯语提示工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。