构建首个大规模真实图像物体合成数据集,支持复杂场景下物体摆放与融合研究。
ORIDa: Object-centric Real-world Image Composition Dataset
- 采集30,000+张真实场景图,含200个独特物体在多位置、多场景下的呈现
- 提供事实-反事实成对数据,每场景5图(4个含物+1个无物背景)
- 适用于生成模型、图像编辑与场景理解等方向的研究者
物体合成任务——将物体放置并协调于多样视觉场景中——随着生成模型的发展成为计算机视觉的重要课题。然而,现有数据集在多样性与规模上仍不足,难以全面覆盖真实世界场景。我们提出ORIDa(Object-centric Real-world Image Composition Dataset),一个大规模、真实拍摄的数据集,包含超过30,000张图像,涵盖200个独特物体在不同位置和场景中的呈现。ORIDa包含两类数据:事实-反事实对与仅事实场景。事实-反事实对由4张物体位于不同位置的场景图和1张不含物体的背景图组成,每场景共5图;仅事实场景则为单张物体置于特定上下文中的图像,拓展了环境多样性。据我们所知,ORIDa是首个具有如此规模与复杂度的公开可获取的真实图像合成数据集。大量分析与实验表明,该数据集对推进物体合成研究具有重要价值。
原文摘要 · Abstract (English)
Object compositing, the task of placing and harmonizing objects in images of diverse visual scenes, has become an important task in computer vision with the rise of generative models. However, existing datasets lack the diversity and scale required to comprehensively explore real-world scenarios. We introduce ORIDa (Object-centric Real-world Image Composition Dataset), a large-scale, real-captured dataset containing over 30,000 images featuring 200 unique objects, each of which is presented across varied positions and scenes. ORIDa has two types of data: factual-counterfactual sets and factual-only scenes. The factual-counterfactual sets consist of four factual images showing an object in different positions within a scene and a single counterfactual (or background) image of the scene without the object, resulting in five images per scene. The factual-only scenes include a single image containing an object in a specific context, expanding the variety of environments. To our knowledge, ORIDa is the first publicly available dataset with its scale and complexity for real-world image composition. Extensive analysis and experiments highlight the value of ORIDa as a resource for advancing further research in object compositing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。