综述越南海量文档识别挑战与大模型新方向
A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions
- 梳理越南文复杂拼写与数据稀缺下的识别难题
- 指出大模型显著提升文本理解但仍有泛化瓶颈
- 适合关注多模态与低资源语言研究者参考
越南语文档分析与识别(DAR)在数字化、信息检索和自动化中具有重要意义。尽管光学字符识别(OCR)与自然语言处理(NLP)取得进展,越南语文本识别仍面临独特挑战:复杂的变音符号、声调变化以及大规模标注数据集的缺乏。传统OCR方法难以应对真实文档的多样性,深度学习虽有潜力,但受限于数据稀缺与泛化能力不足。近期,大语言模型(LLMs)与视觉-语言模型在文本识别与文档理解方面表现突出,为越南语DAR提供新路径。然而,领域自适应、多模态学习与计算效率等挑战依然存在。本文全面回顾现有技术,揭示关键局限,并探讨大模型如何推动该领域变革。我们提出未来方向:数据集建设、模型优化及多模态融合,以提升文档智能水平,促进社区协作创新。
原文摘要 · Abstract (English)
Vietnamese document analysis and recognition (DAR) is a crucial field with applications in digitization, information retrieval, and automation. Despite advancements in OCR and NLP, Vietnamese text recognition faces unique challenges due to its complex diacritics, tonal variations, and lack of large-scale annotated datasets. Traditional OCR methods often struggle with real-world document variations, while deep learning approaches have shown promise but remain limited by data scarcity and generalization issues. Recently, large language models (LLMs) and vision-language models have demonstrated remarkable improvements in text recognition and document understanding, offering a new direction for Vietnamese DAR. However, challenges such as domain adaptation, multimodal learning, and computational efficiency persist. This survey provide a comprehensive review of existing techniques in Vietnamese document recognition, highlights key limitations, and explores how LLMs can revolutionize the field. We discuss future research directions, including dataset development, model optimization, and the integration of multimodal approaches for improved document intelligence. By addressing these gaps, we aim to foster advancements in Vietnamese DAR and encourage community-driven solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。