用隐式神经表示实现跨模态统一风格迁移,支持图文双引导。
OmniStyle-INR: Universal and Multimodal Style Transfer for INRs

- 基于隐式神经表示构建通用视觉风格迁移框架
- 支持图像、视频、3D场景等多模态高质量风格迁移
- 可同时接受文本描述和参考图像进行风格控制
风格迁移是跨多种数据模态的基础且重要的任务,支持基于参考图像和文本描述的创造性调控。尽管高斯点阵(Gaussian Splatting)近期被提出作为2D图像、视频、3D场景及4D动态的统一表示,但其在密集连续域(如图像和视频)中结构上并不最优,所需高斯数量常接近像素总数,实用性受限。相比之下,隐式神经表示(INR)在各类数据域中更自然且流行,具备显著优势:数据压缩、超分辨率能力以及与生成模型的无缝集成。为此,我们提出OmniStyle-INR,一个利用网络化连续表示实现真正通用域的新型框架。该方法成功在所有视觉模态上完成高质量风格迁移,并能无缝结合文本提示与视觉样本进行引导。
原文摘要 · Abstract (English)
Style transfer remains a fundamental and highly important task across various data modalities, enabling creative manipulation conditioned by both reference images and textual descriptions. Recently, methods utilizing Gaussian Splatting have emerged as a unified representation for 2D images, video, 3D scenes, and 4D dynamics. However, representing videos and 2D images with Gaussian Splatting is structurally sub-optimal for dense continuous domains. The number of required Gaussians often approaches the total number of pixels, raising questions about the actual utility of such a representation for these specific modalities. In contrast, Implicit Neural Representations have established themselves as a much more popular and natural choice across all these data domains. Implicit Neural Representations naturally provide significant advantages, including data compression, inherent capabilities for super resolution, and seamless integration with deep generative models. To this end, we introduce OmniStyle-INR, a novel framework that leverages network-based continuous representations as a truly universal domain. Our approach successfully performs high-quality style transfer across all visual modalities, guided seamlessly by both text prompts and visual exemplars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。