用AI分析对话文本,自动评估心理咨询中的互动质量。
Estimating Quality in Therapeutic Conversations: A Multi-Dimensional Natural Language Processing Framework
- 构建多维度NLP框架,从对话动态、话题一致性和情感等4个方面提取42个特征。
- 经数据增强后,分类准确率达88.9%,AUC达94.6%,显著提升效果。
- 适用于临床反馈系统,未来可扩展至语音、表情等多模态分析。
咨询中来访者与治疗师的互动质量是决定疗效的关键因素。本文提出一种基于文本转录的多维自然语言处理框架,客观分类咨询会话中的互动质量。基于253份动机访谈文本(150份高质量,103份低质量),提取了42个跨四个领域的特征:对话动态、语义相似性(话题一致性)、情感分类和问题检测。使用随机森林(RF)、Cat-Boost和支持向量机(SVM)进行超参数调优,并通过分层5折交叉验证训练,于保留测试集上评估。在未增广的平衡数据上,RF准确率最高(76.7%),SVM AUC最高(85.4%)。经SMOTE-Tomek数据增强后,性能显著提升:RF达到88.9%准确率、90.0% F1分数、94.6% AUC;SVM达81.1%准确率、83.1% F1分数、93.6% AUC。特征贡献分析显示,来访者言语的均值与标准差、以及双方语义相似性为关键贡献因素。该框架在原始与增强数据上均表现稳健,F1与召回率持续提升。当前为纯文本方法,但支持未来融合语音语调、面部表情等多模态信息以实现更全面评估。本研究提供了一种可扩展、数据驱动的疗法互动质量评价方法,可为临床提供实时反馈,提升线上线下咨询质量。
原文摘要 · Abstract (English)
Engagement between client and therapist is a critical determinant of therapeutic success. We propose a multi-dimensional natural language processing (NLP) framework that objectively classifies engagement quality in counseling sessions based on textual transcripts. Using 253 motivational interviewing transcripts (150 high-quality, 103 low-quality), we extracted 42 features across four domains: conversational dynamics, semantic similarity as topic alignment, sentiment classification, and question detection. Classifiers, including Random Forest (RF), Cat-Boost, and Support Vector Machines (SVM), were hyperparameter tuned and trained using a stratified 5-fold cross-validation and evaluated on a holdout test set. On balanced (non-augmented) data, RF achieved the highest classification accuracy (76.7%), and SVM achieved the highest AUC (85.4%). After SMOTE-Tomek augmentation, performance improved significantly: RF achieved up to 88.9% accuracy, 90.0% F1-score, and 94.6% AUC, while SVM reached 81.1% accuracy, 83.1% F1-score, and 93.6% AUC. The augmented data results reflect the potential of the framework in future larger-scale applications. Feature contribution revealed conversational dynamics and semantic similarity between clients and therapists were among the top contributors, led by words uttered by the client (mean and standard deviation). The framework was robust across the original and augmented datasets and demonstrated consistent improvements in F1 scores and recall. While currently text-based, the framework supports future multimodal extensions (e.g., vocal tone, facial affect) for more holistic assessments. This work introduces a scalable, data-driven method for evaluating engagement quality of the therapy session, offering clinicians real-time feedback to enhance the quality of both virtual and in-person therapeutic interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。