arXiv:2601.22455cs.CV2026-01中稿 · IEEE TVCG被引 2

用涂鸦生成纹理编辑,智能理解用户意图

ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction

  • 结合多模态大模型与图像生成,解析涂鸦的语义意图
  • 通过全局生成图提取局部纹理,解决目标位置模糊问题
  • 适合需要自由涂鸦式交互的3D内容创作者

交互式3D模型纹理编辑为创建3D资产提供了新机遇,手绘风格交互最具直观性。然而现有方法主要支持草图轮廓绘制,对粗粒度涂鸦交互的支持仍有限。此外,当前方法常因涂鸦指令的抽象性,导致编辑意图模糊和目标语义位置不明确。为此,我们提出ScribbleSense,结合多模态大语言模型(MLLMs)与图像生成模型,有效应对上述挑战。利用MLLM的视觉能力预测涂鸦背后的编辑意图;在明确语义意图后,通过全局生成图像提取局部纹理细节,从而锚定局部语义,缓解目标位置模糊问题。实验表明,该方法充分挖掘了MLLM的优势,在涂鸦式纹理编辑任务中达到当前最优性能。

原文摘要 · Abstract (English)

Interactive 3D model texture editing presents enhanced opportunities for creating 3D assets, with freehand drawing style offering the most intuitive experience. However, existing methods primarily support sketch-based interactions for outlining, while the utilization of coarse-grained scribble-based interaction remains limited. Furthermore, current methodologies often encounter challenges due to the abstract nature of scribble instructions, which can result in ambiguous editing intentions and unclear target semantic locations. To address these issues, we propose ScribbleSense, an editing method that combines multimodal large language models (MLLMs) and image generation models to effectively resolve these challenges. We leverage the visual capabilities of MLLMs to predict the editing intent behind the scribbles. Once the semantic intent of the scribble is discerned, we employ globally generated images to extract local texture details, thereby anchoring local semantics and alleviating ambiguities concerning the target semantic locations. Experimental results indicate that our method effectively leverages the strengths of MLLMs, achieving state-of-the-art interactive editing performance for scribble-based texture editing.

纹理编辑涂鸦交互多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。