用NLP分析德国语料,提升职场倦怠检测的准确性与实用性
Using Natural Language Processing to find Indication for Burnout with Text Classification: From Online Data to Real-World Data
- 构建真实世界文本数据集,包含自由回答与量表反馈
- 发现在线训练模型在真实场景中表现不佳,需优化
- 提供可解释性方案,适合临床与AI研究者协作使用
倦怠被世界卫生组织ICD-11列为一种综合征,源于未能有效管理的长期工作压力,特征为精力耗竭、职业冷漠和效能下降。其流行率因测量方法不一致而差异显著。近年来,自然语言处理(NLP)与机器学习技术为通过文本分析检测倦怠提供了新路径,已有研究展示出高预测准确率。本文在德语文本中推进倦怠检测:(a) 收集匿名的真实世界数据集,包含自由文本回答与旧堡倦怠量表(OLBI)结果;(b) 证明基于GermanBERT的在线数据训练模型在真实场景中的局限性;(c) 提出两个经人工标注的BurnoutExpressions数据集版本,所训练模型在实际应用中表现良好;(d) 通过跨学科焦点小组,提供对用于倦怠检测的AI模型可解释性的定性见解。研究强调,需加强人工智能研究者与临床专家的合作以优化倦怠检测模型。同时,更多真实世界数据对于验证和提升当前基于网络爬取数据的NLP方法至关重要,这些方法通常未在真实环境中评估,难以直接应用于实践。
原文摘要 · Abstract (English)
Burnout, classified as a syndrome in the ICD-11, arises from chronic workplace stress that has not been effectively managed. It is characterized by exhaustion, cynicism, and reduced professional efficacy, and estimates of its prevalence vary significantly due to inconsistent measurement methods. Recent advancements in Natural Language Processing (NLP) and machine learning offer promising tools for detecting burnout through textual data analysis, with studies demonstrating high predictive accuracy. This paper contributes to burnout detection in German texts by: (a) collecting an anonymous real-world dataset including free-text answers and Oldenburg Burnout Inventory (OLBI) responses; (b) demonstrating the limitations of a GermanBERT-based classifier trained on online data; (c) presenting two versions of a curated BurnoutExpressions dataset, which yielded models that perform well in real-world applications; and (d) providing qualitative insights from an interdisciplinary focus group on the interpretability of AI models used for burnout detection. Our findings emphasize the need for greater collaboration between AI researchers and clinical experts to refine burnout detection models. Additionally, more real-world data is essential to validate and enhance the effectiveness of current AI methods developed in NLP research, which are often based on data automatically scraped from online sources and not evaluated in a real-world context. This is essential for ensuring AI tools are well suited for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。