arXiv:2603.25105cs.CL2026-03被引 1

针对心理健康的LLM训练与对话评估难题,提出oMind框架与数据集。

OMIND: Framework for Knowledge Grounded Finetuning and Multi-Turn Dialogue Benchmark for Mental Health LLMs

  • 基于结构化知识检索与LLM筛选生成16.4万条高质量多任务微调数据
  • 构建多轮对话评测集oMind-Chat,含逐轮与整体评分标准
  • 在推理与对话能力上显著超越基线,最高胜率达80%

大语言模型在复杂任务中表现卓越,但在心理健康领域面临三大挑战:高质量可解释、知识驱动的训练数据稀缺;训练范式局限于核心能力;多轮对话评估体系缺失。为此,我们提出oMind框架,包含支持多样化能力的模型训练与对齐机制,并构建了约16.4万条多任务SFT数据集,通过结构化知识检索、LLM辅助筛选及人工审核流程生成。同时,我们推出oMind-Chat——首个专家标注的多轮对话评测数据集,包含逐轮与整体层面的评分标准。在核心能力与对话任务上的实验表明,oMind模型持续优于基线;oMind-LLM在推理能力上提升显著,最高胜率达80%。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities for complex tasks, yet adaptation in medical domain, specifically mental health, poses specific challenges. Mental health is a rising concern globally with LLMs having large potential to help address the same. We highlight three primary challenges for LLMs in mental health - lack of high quality interpretable and knowledge grounded training data; training paradigms restricted to core capabilities, and evaluation of multi turn dialogue settings. Addressing it, we present oMind framework which includes training and aligning LLM agents for diverse capabilities including conversations; high quality ~164k multi-task SFT dataset, as a result of our generation pipeline based on Structured Knowledge retrieval, LLM based pruning, and review actions. We also introduce oMind-Chat - a novel multi turn benchmark dataset with expert annotated turn level and conversation level rubrics. Our diverse experiments on both core capabilities and conversations shows oMind LLMs consistently outperform baselines. oMind-LLM also shows significantly better reasoning with up to 80% win rate.

心理健康对话系统知识增强评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。