用知识增强视觉模型,让AI玩瑞米库布游戏更快更准
Enhancing Computer Vision with Knowledge: a Rummikub Case Study
- 引入外部知识与推理模块,弥补神经网络整体理解短板
- 知识等价于三分之二训练数据,训练时间减半
- 适合需要推理与常识的视觉任务研究者
人工神经网络擅长识别图像中的单个元素,但默认情况下无法正确整合和理解这些元素的整体意义。为缓解这一缺陷,可向网络中加入显式知识与独立推理组件。本文评估了该方法在解决流行桌面游戏瑞米库布(Rummikub)中的应用效果。实验表明,对于此特定任务,额外引入的背景知识价值相当于三分之二的数据集规模,且使训练时间缩短至原时长的一半。
原文摘要 · Abstract (English)
Artificial Neural Networks excel at identifying individual components in an image. However, out-of-the-box, they do not manage to correctly integrate and interpret these components as a whole. One way to alleviate this weakness is to expand the network with explicit knowledge and a separate reasoning component. In this paper, we evaluate an approach to this end, applied to the solving of the popular board game Rummikub. We demonstrate that, for this particular example, the added background knowledge is equally valuable as two-thirds of the data set, and allows to bring down the training time to half the original time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。