arXiv:2602.08775cs.CVcs.CG2026-02

用古老梵文算法实现低资源语音驱动头像,零显卡也能实时生成。

VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational Avatars

  • 基于梵文口诀设计符号化发音到口形转换流程
  • 在仅用CPU的环境下实现高同步率与稳定口型动画
  • 适合教育类离线系统,对硬件要求极低

说话头形象在教育科技中日益普及,以增强内容的社会存在感和参与度。然而,现有许多说话头生成(THG)方法依赖于以GPU为中心的神经渲染、大规模训练数据集或高容量扩散模型,限制了其在离线或资源受限学习环境中的部署。本文提出一种确定性且面向CPU的THG框架——符号化吠陀计算(Symbolic Vedic Computation),将语音转为时间对齐的音素流,将音素映射至紧凑的视觉口形集合,并通过受吠陀经文Urdhva Tiryakbhyam启发的符号协同发音机制生成平滑口形轨迹。一个轻量级2D渲染器执行感兴趣区域(ROI)变形与口部合成,并具备稳定性控制,可在主流CPU上实现实时合成。实验评估了纯CPU环境下同步精度、时间稳定性和身份一致性,并与代表性可行基线进行对比。结果表明,在显著降低计算负载和延迟的同时,仍可达到可接受的唇音同步质量,支持在低端硬件上部署实用的教育类虚拟形象。

原文摘要 · Abstract (English)

Talking-head avatars are increasingly adopted in educational technology to deliver content with social presence and improved engagement. However, many recent talking-head generation (THG) methods rely on GPU-centric neural rendering, large training sets, or high-capacity diffusion models, which limits deployment in offline or resource-constrained learning environments. A deterministic and CPU-oriented THG framework is described, termed Symbolic Vedic Computation, that converts speech to a time-aligned phoneme stream, maps phonemes to a compact viseme inventory, and produces smooth viseme trajectories through symbolic coarticulation inspired by Vedic sutra Urdhva Tiryakbhyam. A lightweight 2D renderer performs region-of-interest (ROI) warping and mouth compositing with stabilization to support real-time synthesis on commodity CPUs. Experiments report synchronization accuracy, temporal stability, and identity consistency under CPU-only execution, alongside benchmarking against representative CPU-feasible baselines. Results indicate that acceptable lip-sync quality can be achieved while substantially reducing computational load and latency, supporting practical educational avatars on low-end hardware. GitHub: https://vineetkumarrakesh.github.io/vedicthg

说话头生成低资源符号计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。