解释为何关系学习未普及,指出数据格式与建模方法的脱节
Why Isn't Relational Learning Taking Over the World?
- 聚焦实体间关系建模,而非仅处理像素和文本
- 企业核心数据多为表格、数据库等关系型格式
- 适合关注知识图谱、数据库智能的从业者阅读
人工智能正通过建模像素、文字和语音占据世界。但世界本质上由具有属性和相互关系的实体(如物体、事件)构成,应直接建模这些关系而非其感知或描述。你可能认为只建模文字和图像,是因为有价值的数据都以文本和图像形式存在。然而,几乎所有公司最核心的数据都在表格、数据库等关系格式中,包含产品编号、学号、交易编号等标识符,不能简单当作数值处理。研究此类数据的领域被称为关系学习、统计关系人工智能等。本文解释为何关系学习尚未广泛普及——除少数限定关系场景外——并指出实现其应有地位所需的关键改进。
原文摘要 · Abstract (English)
Artificial intelligence seems to be taking over the world with systems that model pixels, words, and phonemes. The world is arguably made up, not of pixels, words, and phonemes but of entities (objects, things, including events) with properties and relations among them. Surely we should model these, not the perception or description of them. You might suspect that concentrating on modeling words and pixels is because all of the (valuable) data in the world is in terms of text and images. If you look into almost any company you will find their most valuable data is in spreadsheets, databases and other relational formats. These are not the form that are studied in introductory machine learning, but are full of product numbers, student numbers, transaction numbers and other identifiers that can't be interpreted naively as numbers. The field that studies this sort of data has various names including relational learning, statistical relational AI, and many others. This paper explains why relational learning is not taking over the world -- except in a few cases with restricted relations -- and what needs to be done to bring it to it's rightful prominence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。