首个面向东巴文字语义理解的多模态信息抽取数据集。
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
- 构建东巴文图像与中文语义标注的图文对数据集。
- 包含23,530句级和2,539段级高质量图文对。
- 揭示主流模型在零样本下提取复杂语义仍存挑战。
东巴象形文字是当今世界上唯一仍在使用的象形文字,其图像表意特征蕴含丰富的文化与语境信息。然而,由于缺乏相关数据集,东巴文语义理解研究进展缓慢。为此,我们构建了首个聚焦东巴文字多模态信息抽取的基准数据集——DongbaMIE。该数据集包含23,530句级和2,539段级高质量文本-图像对,图像为东巴文字符,对应中文语义标注涵盖对象、动作、关系和属性四个语义维度。对主流多模态大模型的系统评估表明,在零样本和小样本学习条件下,模型难以高效完成东巴文的信息抽取;尽管监督微调可提升性能,但准确提取复杂语义仍是当前重大挑战。
原文摘要 · Abstract (English)
Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic understanding of Dongba hieroglyphs has progressed slowly. To this end, we constructed \textbf{DongbaMIE} - the first dataset focusing on multimodal information extraction of Dongba pictographs. The dataset consists of images of Dongba hieroglyphic characters and their corresponding semantic annotations in Chinese. It contains 23,530 sentence-level and 2,539 paragraph-level high-quality text-image pairs. The annotations cover four semantic dimensions: object, action, relation and attribute. Systematic evaluation of mainstream multimodal large language models shows that the models are difficult to perform information extraction of Dongba hieroglyphs efficiently under zero-shot and few-shot learning. Although supervised fine-tuning can improve the performance, accurate extraction of complex semantics is still a great challenge at present.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。