针对越南语图文描述生成,提出融合语音结构的多模态融合框架
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

- 构建异构图结构,显式引入越南语语言特征进行多模态融合
- 提出首个大规模越南语图文描述数据集,含1.57万张图像和7.5万条标注
- 解决声调符号混淆问题,52.8%词汇因拼写错误易产生语义歧义
场景文本图像描述需融合视觉特征、OCR识别文本和语言知识三类信息,以生成准确包含图像中可见文字的描述。现有融合方法将文本视为与语言无关,对越南语不适用:越南语为声调语言,变音符号改变词义,OCR错误普遍,词边界模糊。我们主张越南语场景文本描述需采用‘语言感知的多模态融合’,即在融合机制中显式融入语言特有结构知识。基于此,提出通用图融合框架HSTFG,结合学习的空间注意力偏置,并通过拓扑分析发现跨模态图边对图文融合有害。在此基础上,设计专用于越南语语言推理的图级融合模型PhonoSTFG。为支持评估,构建首个大规模越南语场景文本描述数据集ViTextCaps(15,729张图像,74,970条标注),经语言学分析显示52.8%词汇存在变音符号冲突风险。
原文摘要 · Abstract (English)
Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate descriptions that faithfully integrate text visible in images. Existing fusion approaches treat text as language-agnostic, which fails for Vietnamese: a tonal language where diacritics alter word meaning, OCR errors are pervasive, and word boundaries are ambiguous. We argue that Vietnamese scene-text captioning demands \textit{linguistically informed multimodal fusion}, where language-specific structural knowledge is explicitly incorporated into the fusion mechanism. Motivated from these insights, we propose \textbf{HSTFG} (Heterogeneous Scene-Text Fusion Graph), a general-purpose graph fusion framework with learned spatial attention bias, and show through topology analysis that cross-modal graph edges are harmful for scene-text fusion. Building on this finding, we design \textbf{PhonoSTFG} (Phonological Scene-Text Fusion Graph) which specializes graph-level fusion for Vietnamese linguistic reasoning. To support evaluation, we introduce \textbf{ViTextCaps}, the first large-scale Vietnamese scene-text captioning dataset (\textbf{15{,}729} images with \textbf{74{,}970} captions), with comprehensive linguistic analysis showing that 52.8\% of the vocabulary is at risk of diacritic collision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。