arXiv:2507.00263cs.CVcs.LG2025-07

自动识别民宿图片中的房间类型和床位配置,帮游客看清空间布局。

Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections

  • 用视觉+元数据联合建模,分步完成房间分类与聚类。
  • 在少量样本下仍保持高准确率,适合实时处理场景。
  • 特别适合旅游平台、房源智能管理等应用。

度假租赁(VR)平台的快速发展带来了大量未结构化上传的房源图片,缺乏分类导致游客难以理解房屋的空间布局,尤其当存在多个相同类型的房间时。为此,我们提出一种高效机器学习流程,解决房间场景发现与分组问题,并识别每个卧室组的床型。该流程结合监督式房间类型检测模型、重叠相似度检测模型及聚类算法,利用图像间相似性将同一空间的图片归为一组;同时基于多模态大语言模型(MLLM),将卧室组与房源元数据中的床型信息进行视觉匹配。我们分别评估各模块性能,并整体测试该流程,结果显著优于对比方法(如对比学习、预训练嵌入聚类),展现出低延迟、样本高效的优势,适用于实时与数据稀缺环境。

原文摘要 · Abstract (English)

The rapid growth of vacation rental (VR) platforms has led to an increasing volume of property images, often uploaded without structured categorization. This lack of organization poses significant challenges for travelers attempting to understand the spatial layout of a property, particularly when multiple rooms of the same type are present. To address this issue, we introduce an effective approach for solving the room scene discovery and grouping problem, as well as identifying bed types within each bedroom group. This grouping is valuable for travelers to comprehend the spatial organization, layout, and the sleeping configuration of the property. We propose a computationally efficient machine learning pipeline characterized by low latency and the ability to perform effectively with sample-efficient learning, making it well-suited for real-time and data-scarce environments. The pipeline integrates a supervised room-type detection model, a supervised overlap detection model to identify the overlap similarity between two images, and a clustering algorithm to group the images of the same space together using the similarity scores. Additionally, the pipeline maps each bedroom group to the corresponding bed types specified in the property's metadata, based on the visual content present in the group's images using a Multi-modal Large Language Model (MLLM) model. We evaluate the aforementioned models individually and also assess the pipeline in its entirety, observing strong performance that significantly outperforms established approaches such as contrastive learning and clustering with pretrained embeddings.

图像理解空间布局多模态旅游应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。