首个韩语青少年认知扭曲数据集,用多模型协商提升标注质量。
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
- 通过角色轮换的多大模型协商,迭代优化认知扭曲分类。
- 构建10.8万条韩语青少年认知扭曲数据,含合成数据增强多样性。
- 适合心理健康、NLP与大模型应用研究者使用。
认知扭曲是导致青少年抑郁和焦虑等心理问题的负面思维模式。以往基于自然语言处理的研究多集中于小规模成人数据集,对青少年群体关注不足。本研究提出KoACD,首个大规模韩语青少年认知扭曲数据集,包含108,717个实例。采用多大语言模型(LLM)协商方法,通过模型间迭代反馈与角色轮换,降低偏见并提升标签一致性。同时,使用两种方法生成合成数据:认知澄清以提高文本清晰度,认知平衡以增强扭曲类型的多样性表示。通过大模型与专家评估验证,尽管大模型在识别显性标记的认知扭曲上表现良好,但在依赖上下文推理的任务中准确率低于人类评估者。KoACD旨在推动未来认知扭曲检测研究。数据集及实现细节已公开。
原文摘要 · Abstract (English)
Cognitive distortion refers to negative thinking patterns that can lead to mental health issues like depression and anxiety in adolescents. Previous studies using natural language processing (NLP) have focused mainly on small-scale adult datasets, with limited research on adolescents. This study introduces KoACD, the first large-scale dataset of cognitive distortions in Korean adolescents, containing 108,717 instances. We applied a multi-Large Language Model (LLM) negotiation method to refine distortion classification, enabling iterative feedback and role-switching between models to reduce bias and improve label consistency. In addition, we generated synthetic data using two approaches: cognitive clarification for textual clarity and cognitive balancing for diverse distortion representation. Validation through LLMs and expert evaluations showed that while LLMs classified distortions with explicit markers, they struggled with context-dependent reasoning, where human evaluators demonstrated higher accuracy. KoACD aims to enhance future research on cognitive distortion detection. The dataset and implementation details are publicly accessible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。