评估现有语音数据集对协作问题求解模型训练的适用性
An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving
- 基于认知、社交、情绪指标分析语音数据集
- 发现多数现有数据集不满足三人以上团队协作需求
- 提出未来数据集需包含自然对话与多角色互动
本报告评估了现有语音数据集在训练机器学习模型、决策方法和分析算法以提升协作问题求解能力方面的适用性,并列出了未来数据集应具备的要求。问题求解假定由约三至四人组成的团队通过口头交流完成,数据集包含此类团队的语音记录。分析方法基于衡量认知、社会及情感活动与情境的指标。报告还分析了一组广泛用于口语理解研究的数据集,该领域与协作问题求解具有一定相似性。
原文摘要 · Abstract (English)
This report characterized the suitability of existing datasets for devising new Machine Learning models, decision making methods, and analysis algorithms to improve Collaborative Problem Solving and then enumerated requirements for future datasets to be devised. Problem solving was assumed to be performed in teams of about three, four members, which talked to each other. A dataset consists of the speech recordings of such teams. The characterization methodology was based on metrics that capture cognitive, social, and emotional activities and situations. The report presented the analysis of a large group of datasets developed for Spoken Language Understanding, a research area with some similarity to Collaborative Problem Solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。