用可微物理笔模型统一在线与离线手写生成,兼顾结构与真实外观。
Bridging Online and Offline Handwriting via Differentiable Physical Rendering

- 构建可微笔刷模型,将笔迹轨迹映射为像素级视觉特征。
- 在真实机器人书法中实现高保真结构与纹理生成。
- 适合字体设计、生物识别等需兼顾动态与视觉的场景。
真实手写文本生成在字体设计、生物特征认证和机器人书法等领域具有重要意义。现有方法分为两类:在线方法捕捉笔迹动态但缺乏细节纹理,离线方法生成逼真图像却丢失笔画顺序。二者融合面临两大挑战:一是缺乏将笔迹运动学映射到像素外观的显式物理模型;二是缺少成对的轨迹-图像数据集。此外,端到端学习需要跨运动与外观域的可微渲染。为此,我们提出一种紧凑的物理笔刷模型,结合可微渲染模块,将笔迹轨迹转化为风格化图像。所提框架包含四个核心模块:1)文本到笔迹生成器,根据文本与风格图预测目标笔迹;2)笔刷参数观测器,从风格参考中提取笔刷参数;3)可微笔刷渲染器,将笔迹序列与笔刷参数映射为手写图像;4)零样本图像精修模块,通过扩散模型优化渲染结果。大量实验与真实机器人书法演示验证了该方法的有效性,实现了结构与视觉双重保真。
原文摘要 · Abstract (English)
Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calligraphy. Existing methods are typically divided into two independent paradigms: online approaches that estimate handwriting trajectories and offline approaches that synthesize realistic handwriting images. While online models capture structural and temporal dynamics, they often lack fine-grained textures, whereas offline models reproduce realistic appearance but discard stroke order. However, unifying online and offline models remains challenging due to (1) the lack of an explicit physical model linking stroke kinematics to pixel-level appearance and (2) the absence of paired trajectory-image datasets. Moreover, enabling end-to-end learning requires a differentiable rendering process across motion and appearance domains. To address these challenges, we propose a compact physical brush model that bridges stroke dynamics and visual appearance, together with a differentiable rendering module that converts stroke trajectories into stylized images. By integrating these components, we propose a unified online-offline handwriting generation framework via differentiable brush rendering. The proposed framework consists of four core modules: 1) a text-to-stroke generator that predicts the target stroke conditioned on the given text and style image, 2) a brush parameter observer that extracts brush model parameters from style references, 3) a differentiable brush renderer that maps a stroke sequence and physical brush parameters into a handwritten image, and 4) a zero-shot image refiner that refines rendered images via diffusion models. Extensive experiments and real-world robotic calligraphy demonstrations validate our approach, achieving both structural and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。