用动态神经场模拟语音规划,让计算机更像人一样说话
PyPhonPlan: Simulating phonetic planning with dynamic neural fields and task dynamics
- 用动态神经场建模语音规划过程,结合记忆与感知反馈
- 能模拟发音与感知的交互循环,生成符合时间规律的语音轨迹
- 开源工具包,适合语音认知与计算语言学研究者使用
我们提出了PyPhonPlan,一个用于实现基于耦合动态神经场和任务动态仿真的语音规划动力学模型的Python工具包。该工具包提供模块化组件,用于定义规划、感知与记忆场,以及场间耦合、动作输入,并利用场激活轮廓求解可变轨迹问题。通过一个示例应用——耦合记忆场的发音/感知循环模拟,展示了该框架在建模具有时间一致性、神经基础和音位丰富性的交互式语音动态方面的潜力。PyPhonPlan以开源形式发布,包含可执行示例,旨在促进语音通信研究的可复现性、可扩展性和累积性计算发展。
原文摘要 · Abstract (English)
We introduce PyPhonPlan, a Python toolkit for implementing dynamical models of phonetic planning using coupled dynamic neural fields and task dynamic simulations. The toolkit provides modular components for defining planning, perception and memory fields, as well as between-field coupling, gestural inputs, and using field activation profiles to solve tract variable trajectories. We illustrate the toolkit's capabilities through an example application: simulating production/perception loops with a coupled memory field, which demonstrates the framework's ability to model interactive speech dynamics using representations that are temporally-principled, neurally-grounded, and phonetically-rich. PyPhonPlan is released as open-source software and contains executable examples to promote reproducibility, extensibility, and cumulative computational development for speech communication research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。