arXiv:2503.03797cs.SDcs.AI2025-03被引 3

用分组相对策略优化提升语音病理检测准确率

VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection

  • 融合专家混合的Transformer架构,结合分组相对策略优化训练
  • 在合成数据集上准确率、F1和AUC均显著优于传统方法
  • 适合医疗语音分析与自动化诊断系统开发者参考

本研究提出一种新型AI技术——基于分组相对策略优化(GRPO)的专家混合变压器,用于语音健康护理中的语音病理检测。通过架构创新,采用受强化学习启发的先进训练范式,包括近端策略优化(PPO)和分组正则化策略优化(GRPO),以增强模型稳定性和性能。在合成生成的语音病理数据集上的实验表明,所提模型在诊断准确率、F1分数和ROC-AUC方面均显著优于传统方法。这些结果凸显了将变压器架构与新型训练策略结合,在推进自动化语音病理检测方面的潜力,最终有助于提升医疗服务质量。训练与评估代码已开源:https://github.com/enkhtogtokh/voicegrpo

原文摘要 · Abstract (English)

This research introduces a novel AI techniques as Mixture-of-Experts Transformers with Group Relative Policy Optimization (GRPO) for voice health care applications on voice pathology detection. With the architectural innovations, we adopt advanced training paradigms inspired by reinforcement learning, namely Proximal Policy Optimization (PPO) and Group-wise Regularized Policy Optimization (GRPO), to enhance model stability and performance. Experiments conducted on a synthetically generated voice pathology dataset demonstrate that our proposed models significantly improve diagnostic accuracy, F1 score, and ROC-AUC compared to conventional approaches. These findings underscore the potential of integrating transformer architectures with novel training strategies to advance automated voice pathology detection and ultimately contribute to more effective healthcare delivery. The code we used to train and evaluate our models is available at https://github.com/enkhtogtokh/voicegrpo

语音病理Transformer强化学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。