arXiv:2503.17018cs.SDcs.AI2025-03

用符号决策树实现高精度、可解释的音频分类。

Symbolic Audio Classification via Modal Decision Tree Learning

  • 采用符号化决策树方法,替代传统黑箱神经网络。
  • 在年龄、性别、情绪和呼吸疾病诊断任务中均达高准确率。
  • 规则简单透明,适合医疗等需可解释性的场景。

声学分析应用广泛,其中声音分类是近年机器学习关注的重点。现有主流方法多为非符号化,通常基于神经网络,虽性能优异但缺乏透明性。本文针对年龄与性别识别、情绪分类及呼吸系统疾病诊断等音频任务,采用符号化方法——(模态)决策树学习。实验表明,这些任务可通过统一的符号化流程解决,生成的规则简单且准确率高、复杂度低。理论上,此类系统可集成至自主对话系统,适用于医院或诊所的自动预约代理等场景。

原文摘要 · Abstract (English)

The range of potential applications of acoustic analysis is wide. Classification of sounds, in particular, is a typical machine learning task that received a lot of attention in recent years. The most common approaches to sound classification are sub-symbolic, typically based on neural networks, and result in black-box models with high performances but very low transparency. In this work, we consider several audio tasks, namely, age and gender recognition, emotion classification, and respiratory disease diagnosis, and we approach them with a symbolic technique, that is, (modal) decision tree learning. We prove that such tasks can be solved using the same symbolic pipeline, that allows to extract simple rules with very high accuracy and low complexity. In principle, all such tasks could be associated to an autonomous conversation system, which could be useful in different contexts, such as an automatic reservation agent for an hospital or a clinic.

音频分类决策树可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。