arXiv:2509.14601cs.DBcs.AI2025-09被引 3

让文字图像等无结构数据也能高效计算,突破传统系统局限。

A Case for Computing on Unstructured Data

  • 构建三阶段双向流程:提取隐含结构、处理结构、还原无结构格式
  • 支持对文本图像音频视频的分析计算,保持原始表达丰富性
  • 适合需要处理真实世界多模态数据的研究与应用开发者

文本、图像、音频和视频等无结构数据构成了全球信息的主体,但传统数据系统依赖结构化格式进行计算,难以有效支持。本文提出一种新范式——无结构数据计算,包含三个阶段:提取隐含结构、通过数据处理技术转换结构、再投影回无结构格式。该双向流程使无结构数据可利用结构化计算的分析能力,同时保留人类与AI可读的原始表现形式。文章通过两个应用场景说明该范式,并提出需在新数据系统MXFlow中开发的关键研究组件。

原文摘要 · Abstract (English)

Unstructured data, such as text, images, audio, and video, comprises the vast majority of the world's information, yet it remains poorly supported by traditional data systems that rely on structured formats for computation. We argue for a new paradigm, which we call computing on unstructured data, built around three stages: extraction of latent structure, transformation of this structure through data processing techniques, and projection back into unstructured formats. This bi-directional pipeline allows unstructured data to benefit from the analytical power of structured computation, while preserving the richness and accessibility of unstructured representations for human and AI consumption. We illustrate this paradigm through two use cases and present the research components that need to be developed in a new data system called MXFlow.

无结构数据数据系统多模态计算范式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。