用临床分级框架评估自杀风险,发现平台仅靠标记无法准确判断严重程度。
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

- 基于临床量表构建四等级风险标注数据集,评估不同方法的分级能力。
- 主流审核接口能区分高低风险(高风险F1 0.860),但严重程度判断差(宏平均F1 0.395)。
- 专家设计的提示词比微调更有效,尤其适用于简洁的临床语境。
当前内容审核系统旨在识别违规内容,而非衡量渐进式临床风险。然而,平台责任不仅限于检测:对被动表达痛苦与主动计划并具备手段获取者,所需响应差异显著,加州参议院法案243等新规已将此区分纳入合规要求。本文探讨现有安全信号在恢复临床有意义风险等级方面的表现。我们发布了一个包含516条r/SuicideWatch帖子的基准数据集,由持证精神科医生依据哥伦比亚自杀严重度量表(Columbia Suicide Severity Rating Scale)的四等级序数标注体系(指示、意念、行为、尝试)进行评级,并在七种序数感知指标下评估了厂商审核接口、提示式LLM及监督基线模型。三项发现:厂商审核接口能较好分离低/高风险帖子(高风险F1 0.860),但严重程度判别能力弱(宏平均F1 0.395),系统性高估最严重类别;基于临床知识的零样本提示可大幅缩小差距(宏平均F1 0.562),而专家撰写的语境提示(非微调、推理增强或朴素多代理聚合)是关键杠杆。推理价值取决于语体:在冗长嘈杂的Reddit帖文中起反作用,在简短临床文本中则显著提升表现。我们认为,比例性的关怀义务应指向分级风险评估,而非二元标记,并公开评估框架以支持该方向。
原文摘要 · Abstract (English)
Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emerging regulation (e.g., California Senate Bill 243) is turning that distinction into a compliance requirement. We therefore ask how well deployed safety signals recover clinically meaningful severity. We release a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluate moderation APIs, prompted LLMs, and supervised baselines under seven ordinal-aware metrics. Three findings. Vendor moderation APIs separate low- from high-severity posts well (0.860 high-risk F1) but measure severity poorly (0.395 macro F1), systematically over-predicting the most severe category. Clinically grounded zero-shot prompting recovers much of that gap (0.562 macro F1), and expert-authored framing (not fine-tuning, added reasoning, or naive multi-agent aggregation) is the effective lever. The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements. We argue graded severity, not a binary flag, is what a proportionate duty of care requires, and release our evaluation framework to support that measurement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。