arXiv:2605.18680cs.CV2026-05中稿 · CVPR

用3D概念框架解决虚拟形象生成中的文本歧义问题

CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation

论文配图:CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation
图 1 · 摘自论文原文
  • 通过3D概念骨架明确用户意图的全局风格与空间关系
  • 在复杂提示下实现92.3%的组合正确率,优于基线模型
  • 适合需要高精度形象生成的元宇宙创作者使用

元宇宙平台依赖创作者驱动的市场,用户从离散的、按分类标签标记的3D资产(如上衣、下装、鞋子、配饰)中组装虚拟形象,受严格类别与拓扑约束。尽管用户日益期望自由文本控制,但纯文本检索存在缺陷:自然语言与平台分类体系不一致,元数据常含噪声或非正式信息,独立检索的部件常出现风格不一致或几何不兼容。本文提出CMAG——一种基于概念的检索与验证合成框架。给定提示后,CMAG首先生成中间3D概念骨架,通过全局空间与风格上下文消除文本歧义;并行地,视图感知部件发现模块通过提示分解与文本引导分割提取局部视觉证据。提示条件分类路由模块确保类别覆盖并解决语义到分类的错位,随后混合式类别内检索器结合部件融合与概念残差回退(特征抑制)。最终,智能体式视觉-语言模型跨类别筛选与重排序候选,并驱动迭代验证循环,从目录资产中组装符合提示、拓扑一致的虚拟形象。我们在多样组合提示上评估了CMAG,结果表明其在检索鲁棒性与组合正确性方面显著优于强基线,凸显3D概念骨架在提示模糊情况下的重要性。

原文摘要 · Abstract (English)

Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoes, accessories) under strict category and topology constraints. While users increasingly expect free-form text control, text-only retrieval is brittle: natural language is ambiguous with respect to platform taxonomies, metadata is often noisy or informal, and independently retrieved components can be stylistically inconsistent or geometrically incompatible. We propose \textbf{CMAG}, a concept-scaffolded retrieval and verified composition framework for marketplace avatar generation. Given a prompt, CMAG first synthesizes an intermediate 3D concept scaffold that disambiguates intent beyond text by providing global spatial and stylistic context. In parallel, a view-aware part discovery module extracts localized visual evidence via prompt decomposition and text-grounded segmentation. A prompt-conditioned taxonomy router enforces category coverage and resolves semantic-to-taxonomic mismatch, after which a hybrid category-wise retriever combines part-based fusion with a concept-residual fallback using feature suppression. Finally, an agentic vision--language model filters and re-ranks candidates across categories and drives an iterative verification loop to assemble prompt-faithful, topologically consistent avatars from catalog assets. We evaluate CMAG on diverse compositional prompts and demonstrate improved retrieval robustness and compositional correctness compared to strong baselines, highlighting the importance of 3D concept scaffolding under prompt ambiguity.

虚拟形象概念骨架多模态检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。