针对营销电话中购买意愿分类,提出多段多任务融合网络
MSMT-FN: Multi-segment Multi-task Fusion Network for Marketing Audio Classification
- 将音频分段后并行处理,融合多任务学习提升分类精度
- 在自建数据集和多个基准上性能超越或持平当前最佳方法
- 适合语音情感分析与商业智能领域研究者参考使用
音频分类在情感分析与情绪识别中具有关键作用,尤其适用于分析营销电话中的客户态度。从大量音频数据中高效分类客户购买意愿仍是挑战。本文提出一种专为该商业需求设计的多段多任务融合网络(MSMT-FN)。在自研的MarketCalls数据集以及经典基准(CMU-MOSI、CMU-MOSEI、MELD)上的评估显示,MSMT-FN始终优于或匹配现有最先进方法。此外,我们新构建的MarketCalls数据集可应需提供,代码已开源至GitHub仓库MSMT-FN,以促进音频分类领域的进一步研究与发展。
原文摘要 · Abstract (English)
Audio classification plays an essential role in sentiment analysis and emotion recognition, especially for analyzing customer attitudes in marketing phone calls. Efficiently categorizing customer purchasing propensity from large volumes of audio data remains challenging. In this work, we propose a novel Multi-Segment Multi-Task Fusion Network (MSMT-FN) that is uniquely designed for addressing this business demand. Evaluations conducted on our proprietary MarketCalls dataset, as well as established benchmarks (CMU-MOSI, CMU-MOSEI, and MELD), show MSMT-FN consistently outperforms or matches state-of-the-art methods. Additionally, our newly curated MarketCalls dataset will be available upon request, and the code base is made accessible at GitHub Repository MSMT-FN, to facilitate further research and advancements in audio classification domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。