对比LoRA/QLoRA与偏好优化,指导心理文本分类的训练策略选择。
From Baselines to Preferences: A Comparative Study of LoRA/QLoRA and Preference Optimization for Mental Health Text Classification
- 从基线到偏好优化,系统比较不同微调方法的适用场景。
- 偏好优化效果波动大,配置与数据平衡显著影响性能。
- 建议先用透明基线,再选择性使用偏好优化,提升可复现性。
心理文本分类广泛采用现代适配方法,但如何选择优化策略、何时使用及原因仍缺乏实践指导。本文针对联合心理健康分类任务,系统比较从强基线到渐进式专业化技术的优化路径。首先建立经典与编码器基线,再在多种目标与优化设置下评估参数高效监督微调(LoRA/QLoRA),最后考察基于偏好的优化方法(DPO、ORPO、KTO),包括类别均衡训练。不强调单一指标,而是关注方法论洞察:性能如何随目标设定、适配器选择、优化器行为、上下文窗口和类别平衡干预而变化。结果表明,优化效果高度依赖方法:部分方法表现稳定且可迁移,部分则对配置和数据分布敏感。偏好优化尤其在不同目标间差异显著,说明方法选择比简单增加偏好训练阶段更为关键。核心贡献是为心理健康NLP提供清晰的优化叙事:从透明基线出发,进行可控微调,仅在收益明确时选用偏好优化。这构建了一个超越模型架构选择的可复现、实用的训练策略框架。
原文摘要 · Abstract (English)
Mental health text classification has rapidly adopted modern adaptation methods, yet practical guidance on which optimization strategy to use, when, and why remains limited. This paper presents a systematic comparative study of optimization pathways for a joint mental-health classification task, moving from strong vanilla baselines to progressively more specialized techniques. We first establish classical and encoder references, then examine parameter-efficient supervised fine-tuning with LoRA/QLoRA under multiple objective and optimization settings, and finally evaluate preference-based optimization with DPO, ORPO, and KTO, including class-rebalanced training. Rather than emphasizing a single headline score, we focus on methodological insight: how performance changes with objective formulation, adapter choice, optimizer behavior, context windowing, and class-balance intervention. The results show that optimization effects are highly method-dependent: some approaches deliver stable, transferable gains, while others are sensitive to configuration and data balance. Preference optimization, in particular, exhibits large variation across objectives, indicating that method selection is more consequential than simply adding a preference-training stage. The central contribution is a clear optimization narrative for mental health NLP: start from transparent baselines, apply controlled tuning, and use preference optimization selectively where its gains are demonstrable. This provides a reproducible and practically grounded framework for choosing effective training strategies beyond architecture choice alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。