用对话模型检测抑郁,通过提示工程提升评估一致性。
DS@GT at eRisk 2025: From prompts to predictions, benchmarking early depression detection with conversational agent based assessments and temporal attention models
- 设计提示让大模型按BDI-II标准进行结构化评估。
- 在无真实标签下实现0.50的DCHR与0.89的ADODL得分。
- 适合研究心理评估自动化与对话系统应用者阅读。
本文总结了DS@GT团队参与eRisk 2025两项挑战的情况。针对基于大语言模型(LLMs)的对话式抑郁检测试点任务,我们采用提示工程策略,让多种LLM执行基于BDI-II量表的评估,并生成结构化JSON输出。由于缺乏真实标签,我们通过跨模型一致性和内部一致性进行评估。提示设计使模型输出与BDI-II标准对齐,支持对对话线索影响症状预测的分析。最佳提交方案在官方排行榜中位列第二,取得DCHR = 0.50、ADODL = 0.89、ASHR = 0.27的成绩。
原文摘要 · Abstract (English)
This Working Note summarizes the participation of the DS@GT team in two eRisk 2025 challenges. For the Pilot Task on conversational depression detection with large language-models (LLMs), we adopted a prompt-engineering strategy in which diverse LLMs conducted BDI-II-based assessments and produced structured JSON outputs. Because ground-truth labels were unavailable, we evaluated cross-model agreement and internal consistency. Our prompt design methodology aligned model outputs with BDI-II criteria and enabled the analysis of conversational cues that influenced the prediction of symptoms. Our best submission, second on the official leaderboard, achieved DCHR = 0.50, ADODL = 0.89, and ASHR = 0.27.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。