用多模型辩论自动标注心理健康与网络安全数据,提升标注精度。
Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
- 多大模型细粒度辩论,带置信度反馈协同标注
- 在多个任务上提升9.9点宏F1,效果稳定
- 适合需要高精度多标签标注的健康与安全研究
真实世界指标在自然语言处理中至关重要,如心理健康的事件和网络安全部分的危险行为,但因其多标签和动态特性,人工标注成本高。大语言模型虽具自动化标注潜力,但在多标签场景仍具挑战。本文提出置信度感知的细粒度辩论框架(CFD),模拟人类协作标注,通过细粒度交流实现更优的自动多标签扩充。构建两个专家标注资源:心理健康福祉数据集中的生活事件与症状标注,以及新的在线安全‘晒娃’行为数据集。实验表明,CFD在各项任务中表现最稳健,且细粒度置信度的有效性取决于其质量和差异性。进一步评估无需训练的融合策略,发现自动扩充可显著提升下游任务性能,最高达9.9点宏F1,最优融合方式取决于指标与任务目标的相关性。
原文摘要 · Abstract (English)
Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often costly and/or difficult due to its multi-label and dynamic nature. Large Language Models (LLMs) show promising potential for automated annotation, but the multi-label setting remains challenging. In this work, we propose a Confidence-Aware Fine-Grained Debate (CFD) framework that simulates human collaborative annotation using fine-grained communication to better support automated multi-label enrichment. We introduce two expert-annotated resources: life-event and symptom annotation for a mental health well-being dataset, and a new online safety sharenting dataset. Experiments show that CFD achieves the most robust enrichment performance across tasks, and that the benefit of fine-grained confidence is influenced by its quality and variability. We further evaluate training-free strategies for incorporating enrichment indicators into downstream tasks and show that automated enrichment consistently improves performance, by up to 9.9 Macro-F1 points, with the most effective integration format depending on how the indicator relates to the downstream objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。