arXiv:2506.02924cs.CLcs.IR2025-06被引 1

用微调模型和合成数据提升抑郁症状检索效果

INESC-ID @ eRisk 2025: Exploring Fine-Tuned, Similarity-Based, and Prompt-Based Approaches to Depression Symptom Identification

  • 将症状检索转为每个症状的二分类任务,结合微调与相似度匹配
  • 微调模型在验证集上表现最佳,尤其在加入合成数据后改善了类别不平衡
  • 针对不同症状设计差异化策略,集成方法最终胜出16支队伍

本文介绍我们团队在eRisk 2025任务1(抑郁症状搜索)中的方法。给定一组句子和贝克抑郁量表-第二版(BDI-II)问卷,参赛者需提交每种抑郁症状最多1,000句相关句子,并按相关性排序。评估采用信息检索标准指标,包括平均精度(AP)和R-precision(R-PREC)。训练数据仅标注句子是否与某个BDI症状相关,因此我们将问题建模为每个症状的二分类任务,并据此进行评估。我们划分训练与验证集,探索了基础模型微调、句间相似度计算、大语言模型提示(LLM prompting)及集成方法。验证结果显示,微调基础模型表现最优,特别是通过合成数据缓解类别不平衡后。不同症状的最佳方法各异。基于此,我们设计五次独立测试运行,其中两次采用集成策略,最终在官方IR评估中取得最高分,超越16支其他团队。

原文摘要 · Abstract (English)

In this work, we describe our team's approach to eRisk's 2025 Task 1: Search for Symptoms of Depression. Given a set of sentences and the Beck's Depression Inventory - II (BDI) questionnaire, participants were tasked with submitting up to 1,000 sentences per depression symptom in the BDI, sorted by relevance. Participant submissions were evaluated according to standard Information Retrieval (IR) metrics, including Average Precision (AP) and R-Precision (R-PREC). The provided training data, however, consisted of sentences labeled as to whether a given sentence was relevant or not w.r.t. one of BDI's symptoms. Due to this labeling limitation, we framed our development as a binary classification task for each BDI symptom, and evaluated accordingly. To that end, we split the available labeled data into training and validation sets, and explored foundation model fine-tuning, sentence similarity, Large Language Model (LLM) prompting, and ensemble techniques. The validation results revealed that fine-tuning foundation models yielded the best performance, particularly when enhanced with synthetic data to mitigate class imbalance. We also observed that the optimal approach varied by symptom. Based on these insights, we devised five independent test runs, two of which used ensemble methods. These runs achieved the highest scores in the official IR evaluation, outperforming submissions from 16 other teams.

抑郁识别大模型应用信息检索数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。