arXiv:2412.00955cs.CV2024-12中稿 · WACV 2025被引 10

构建首个跨类型跨区域的万能平面图数据集,推动建筑语义理解发展。

WAFFLE: Multimodal Floorplan Understanding in the Wild

  • 从网络数据中收集近2万张多模态平面图,覆盖多样建筑类型与地区。
  • 利用大语言模型和多模态基础模型提取图像与噪声元数据中的语义信息。
  • 支持判别与生成任务,为建筑理解研究提供新基准,适合建筑与AI交叉研究者。

建筑是人类文化的核心组成部分,正日益通过计算方法进行分析。然而,现有建筑理解研究主要聚焦于自然图像,忽视了定义建筑结构的根本要素——平面图。而现有的平面图理解工作范围极为有限,通常仅针对单一语义类别和地区(如某一国家的公寓平面图)。本文提出WAFFLE,一个包含近2万张平面图及其元数据的新型多模态平面图理解数据集,数据源自互联网,涵盖多种建筑类型、地理位置与数据格式。我们采用大语言模型和多模态基础模型,对这些图像及其伴随的噪声元数据进行清洗与语义信息提取。实验表明,WAFFLE使此前无法实现的新颖建筑理解任务(包括判别与生成任务)成为可能。我们将公开发布WAFFLE数据集、代码及训练好的模型,为研究社区提供学习建筑语义的新基础。

原文摘要 · Abstract (English)

Buildings are a central feature of human culture and are increasingly being analyzed with computational methods. However, recent works on computational building understanding have largely focused on natural imagery of buildings, neglecting the fundamental element defining a building's structure -- its floorplan. Conversely, existing works on floorplan understanding are extremely limited in scope, often focusing on floorplans of a single semantic category and region (e.g. floorplans of apartments from a single country). In this work, we introduce WAFFLE, a novel multimodal floorplan understanding dataset of nearly 20K floorplan images and metadata curated from Internet data spanning diverse building types, locations, and data formats. By using a large language model and multimodal foundation models, we curate and extract semantic information from these images and their accompanying noisy metadata. We show that WAFFLE enables progress on new building understanding tasks, both discriminative and generative, which were not feasible using prior datasets. We will publicly release WAFFLE along with our code and trained models, providing the research community with a new foundation for learning the semantics of buildings.

建筑理解多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。