arXiv:2504.21017cs.CLcs.LG2025-04被引 3

首个越南语新冠问答数据集,助力AI防疫研究。

ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese

  • 构建首个面向越南语的新冠机器阅读理解数据集
  • 包含多段落答案抽取,支持复杂问答任务
  • 适合越南语NLP与医疗AI研究者使用

新冠疫情持续两年,全球累计确诊超5.22亿例,死亡超600万例(越南近1000万例,死亡超4.3万例),对经济与社会造成严重冲击。奥密克戎变异株突破各国防控措施,感染人数迅速攀升,治疗与防疫资源普遍超载。在此背景下,人工智能在疫情防控中的应用极为迫切。已有大量研究将AI用于新冠预防,其中机器阅读理解(MRC)也日益重要。为此,我们创建了首个针对越南语的新冠机器阅读理解数据集——ViQA-COVID,可用于构建模型与系统,助力疾病防控。此外,该数据集也是首个越南语多跨度答案抽取型MRC数据集,旨在推动越南语及多语言领域的机器阅读理解研究。

原文摘要 · Abstract (English)

After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million deaths worldwide (including nearly ten million cases and over forty-three thousand deaths in Vietnam). Economy and society are both severely affected. The variant of COVID-19, Omicron, has broken disease prevention measures of countries and rapidly increased number of infections. Resources overloading in treatment and epidemics prevention is happening all over the world. It can be seen that, application of artificial intelligence (AI) to support people at this time is extremely necessary. There have been many studies applying AI to prevent COVID-19 which are extremely useful, and studies on machine reading comprehension (MRC) are also in it. Realizing that, we created the first MRC dataset about COVID-19 for Vietnamese: ViQA-COVID and can be used to build models and systems, contributing to disease prevention. Besides, ViQA-COVID is also the first multi-span extraction MRC dataset for Vietnamese, we hope that it can contribute to promoting MRC studies in Vietnamese and multilingual.

机器阅读理解越南语新冠多跨度抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。