通过反事实推理生成更协调的室内家具布局
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

- 构建动态超图风格场捕捉家具间高阶依赖
- 反事实评估候选家具与整体风格的兼容性
- 适合需要保持布局固定的室内设计场景
固定布局的室内家具搭配需在不改变家具类别、位置、朝向和尺度的前提下,选出协调的家具组合。现有方法通常独立检索家具或依赖静态局部关系,导致组合后出现形状、材质和颜色冲突。本文提出StyleForge,一种基于动态超图风格场的场景级结构化选择框架。冻结的多模态大语言模型从开放风格请求和固定布局中提取结构化风格先验,而StyleForge为每个家具位维护可学习的候选分布。在目标风格条件下,动态超图风格场自适应激活并加权布局诱导的超边,以捕捉家具间的高阶依赖关系。反事实风格偏好学习将每个候选视为当前风格场中的局部替换,使用马氏能量评估其上下文兼容性。训练过程中交替优化风格场与候选逻辑值。推理时模型保持冻结,仅通过测试时训练更新房间特定的候选逻辑值,随着全局场景上下文演变逐步修正跨位风格冲突。在3D-FRONT数据集上的实验表明,该方法在家具检索和场景级风格一致性方面达到当前最优水平,生成的固定布局家具组合比基于物体和场景的检索基线更协调。
原文摘要 · Abstract (English)
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。