为心理支持对话提供可解释的多维度审计工具
CounselReflect: A Toolkit for Auditing Mental-Health Dialogues
- 整合12种任务专用指标与69项文献衍生评分标准
- 生成会话级摘要与逐轮打分,附证据片段支持
- 支持实时使用,适合研究人员与临床专家
心理支持正越来越多地通过对话系统(如基于大模型的工具)实现,但用户缺乏结构化方式来评估所获支持的质量与潜在风险。我们提出CounselReflect,一个端到端的心理支持对话审计工具包。该工具不输出单一模糊评分,而是提供多维度结构化报告,包括会话级摘要、逐轮评分及证据关联的摘录,支持透明审查。系统融合两类评估信号:(i) 12种由任务特定预测器生成的模型指标;(ii) 基于文献的69项评分标准库,以及可配置的大语言模型裁判支持的自定义指标。CounselReflect以网页应用、浏览器插件和命令行界面形式提供,支持实时与大规模使用。人类评估包含20名参与者用户研究和6名心理健康专业人士的专家评审,表明其具备可理解性、可用性和可信度。演示视频与完整源码均已公开。
原文摘要 · Abstract (English)
Mental-health support is increasingly mediated by conversational systems (e.g., LLM-based tools), but users often lack structured ways to audit the quality and potential risks of the support they receive. We introduce CounselReflect, an end-to-end toolkit for auditing mental-health support dialogues. Rather than producing a single opaque quality score, CounselReflect provides structured, multi-dimensional reports with session-level summaries, turn-level scores, and evidence-linked excerpts to support transparent inspection. The system integrates two families of evaluation signals: (i) 12 model-based metrics produced by task-specific predictors, and (ii) rubric-based metrics that extend coverage via a literature-derived library (69 metrics) and user-defined custom metrics, operationalized with configurable LLM judges. CounselReflect is available as a web application, browser extension, and command-line interface (CLI), enabling use in real-time settings as well as at scale. Human evaluation includes a user study with 20 participants and an expert review with 6 mental-health professionals, suggesting that CounselReflect supports understandable, usable, and trustworthy auditing. A demo video and full source code are also provided.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。