PragyaDoc让低资源多语种医疗文档在印度本地化可用
PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings

- 四层架构融合光学识别与几何语义信息
- 支持22种印地语系语言的医疗文档理解
- 适合乡村医生、家庭成员等非专业用户
印度有22种官方语言,但绝大多数医疗文档仅以英语存在,导致农村居民、社区健康工作者(ASHA)及患者家属因语言障碍无法获取关键医疗信息。本文提出普拉吉亚文档(PragyaDoc)——一种通用文档智能框架,通过四层流水线解决该问题:并行集成式OCR提取层、几何-词汇融合层、确定性领域结构化层,以及双大模型医学推理与定位层。该框架可在低资源条件下实现多语言医疗文档的理解,显著提升非英语使用者对医疗信息的可及性。
原文摘要 · Abstract (English)
India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information - rural populations, ASHA workers, and patient families - are functionally excluded from understanding it. This paper presents PragyaDoc, a Universal Document Intelligence Framework that addresses this gap through a four-layer pipeline: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。