arXiv:2603.27055cs.CLcs.IR2026-03中稿 · Publication as a B…

将文本数据融入数据集成,解决多源异构数据融合难题

Text Data Integration

  • 提出将自由文本纳入数据集成框架,突破传统仅处理结构化数据的局限
  • 指出文本数据整合面临语义歧义、格式不一等核心挑战
  • 适合数据工程、知识图谱方向研究者参考

数据形式多样,可分为结构化(如关系表、键值对)和非结构化(如文本、图像)两类。当前机器在处理遵循精确模式的结构化数据方面表现良好,但数据异构性带来了存储与处理的挑战。数据集成作为数据工程的关键环节,旨在整合不同来源的数据并为用户提供统一访问。以往系统主要聚焦结构化数据,而未充分利用蕴含丰富知识的非结构化文本数据。本文首先论证文本数据集成的必要性,随后分析其面临的挑战、现有技术进展及开放问题。

原文摘要 · Abstract (English)

Data comes in many forms. From a shallow perspective, they can be viewed as being either in structured (e.g., as a relation, as key-value pairs) or unstructured (e.g., text, image) formats. So far, machines have been fairly good at processing and reasoning over structured data that follows a precise schema. However, the heterogeneity of data poses a significant challenge on how well diverse categories of data can be meaningfully stored and processed. Data Integration, a crucial part of the data engineering pipeline, addresses this by combining disparate data sources and providing unified data access to end-users. Until now, most data integration systems have leaned on only combining structured data sources. Nevertheless, unstructured data (a.k.a. free text) also contains a plethora of knowledge waiting to be utilized. Thus, in this chapter, we firstly make the case for the integration of textual data, to later present its challenges, state of the art and open problems.

数据集成文本处理异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。