arXiv:2502.17793cs.CVcs.AI2025-02ACL被引 6

让AI设计出功能完整的新概念,还能保持视觉新颖性。

SYNTHIA: Novel Concept Design with Affordance Composition

  • 用分层语义结构拆解设计元素,实现多功能组合
  • 人类评估显示新颖性提升25.1%,功能一致性提升14.7%
  • 适合需要创意原型生成的工业设计与产品开发

文本到图像(T2I)模型推动了人工智能驱动的设计创新。尽管现有研究聚焦于概念的语义与风格变化,功能连贯性——即多个可用性整合为统一概念——仍被忽视。本文提出SYNTHIA框架,基于目标可用性生成新颖且功能一致的设计。该方法利用分层概念本体,将设计分解为部件与可用性,作为功能一致性的关键基础。我们还构建了一种基于本体的课程学习策略,通过对比训练逐步优化T2I模型,使其在保持视觉新颖性的前提下学会组合可用性。具体而言,(i)逐步增加可用性间距离,引导模型从基础关联过渡到融合不同部件的复杂组合;(ii)通过对比目标强制学习表征远离已有概念,确保视觉新颖性。实验表明,SYNTHIA在人类评估中相比现有最优模型,新颖性提升25.1%,功能连贯性提升14.7%。

原文摘要 · Abstract (English)

Text-to-image (T2I) models enable rapid concept design, making them widely used in AI-driven design. While recent studies focus on generating semantic and stylistic variations of given design concepts, functional coherence--the integration of multiple affordances into a single coherent concept--remains largely overlooked. In this paper, we introduce SYNTHIA, a framework for generating novel, functionally coherent designs based on desired affordances. Our approach leverages a hierarchical concept ontology that decomposes concepts into parts and affordances, serving as a crucial building block for functionally coherent design. We also develop a curriculum learning scheme based on our ontology that contrastively fine-tunes T2I models to progressively learn affordance composition while maintaining visual novelty. To elaborate, we (i) gradually increase affordance distance, guiding models from basic concept-affordance association to complex affordance compositions that integrate parts of distinct affordances into a single, coherent form, and (ii) enforce visual novelty by employing contrastive objectives to push learned representations away from existing concepts. Experimental results show that SYNTHIA outperforms state-of-the-art T2I models, demonstrating absolute gains of 25.1% and 14.7% for novelty and functional coherence in human evaluation, respectively.

概念设计功能生成视觉新颖文本生成图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。