FAST-MEL实现高效精准的多模态实体链接,速度更快、存储更省。
FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking
- 用固定大小向量表示实体文本与视觉信息,轻量化编码。
- 精度媲美顶尖系统,速度提升1000倍,存储减少10倍。
- 适合大规模实际应用,尤其注重效率与资源的场景。
多模态实体链接(MEL)旨在将非结构化数据中的文本和视觉实体提及匹配到知识库(KB)中的对应实体。为在大规模实际场景中有效运行,MEL系统需兼顾高准确率、计算效率和存储效率,即构建紧凑高效的KB索引。本文指出,现有先进系统无法同时满足这三项要求。为此,我们提出FAST-MEL,一种基于轻量级编码器的MEL解决方案,采用新颖且紧凑的固定尺寸向量表示每个实体或提及的文本与视觉信息。该方法在准确率上达到最佳系统水平,但推理速度提升三个数量级;存储消耗较最快系统降低一个数量级。
原文摘要 · Abstract (English)
Multimodal entity linking (MEL) is the task that consists of matching textual and visual mentions of entities in unstructured data to their corresponding entities in a knowledge base (KB). To be effective in large-scale practical settings, MEL systems must meet three objectives: high linking accuracy, computational efficiency, and storage efficiency, i.e., a compact yet efficient index of the KB. In this paper, we highlight that state-of-the-art systems fail to simultaneously satisfy these 3 requirements. To meet this three-fold objective, we propose FAST-MEL, a lightweight encoder-based MEL solution that relies on a novel and compact fixed-size vectorized representation of both the textual and visual information of each entity or mention. It matches the accuracy of the best systems but performs three orders of magnitude faster. It also consumes one order of magnitude less storage than the fastest systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。