为孕产期心理支持设计安全可控的多模态AI后端,自动识别风险并阻断不当生成。
A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian
- 分层建模用户情绪、心理社会状态与安全风险,实现渐进式评估
- 融合本地与语音外部证据,高风险时触发安全拦截,召回率达100%
- 适合心理支持系统开发者,关注安全机制与临床落地的团队
本文提出Anian,一种面向孕产期心理支持与正念干预路由的安全受控多模态AI后端。该系统不用于精神疾病诊断或替代临床干预。其模块化流程将生成式AI置于结构化状态表征、保守风险融合与响应门控之后。用户文本或语音转录内容被映射至四个关联层级:L1情绪状态、L2心理社会构念、L3安全风险、L4干预路径。本地文本与规则型安全证据,结合外部语音衍生证据,通过最高风险优先规则融合:S_fusion = max(S_local, S_external)。当融合风险为中或高时,普通生成回应与语音合成被阻断,替换为固定安全内容及人工支持提示。内部原型评估基于约858,295条来自公开情绪、对话、心理健康及中文对话语料库的归一化记录,在弱标签与规则推导框架下进行。微平均F1得分分别为:L1情绪分类0.9604,L2心理社会构念0.9144,L4路由0.9742。在233样本的安全压力测试中,L3规则引擎在预设场景下实现1.0000的高风险召回率。结果支持标签框架与门控逻辑的内部可行性,但未验证临床有效性、诊断准确性、真实世界安全或干预效果。本文报告架构、本体、安全融合机制、原型评估、错误分析计划及专家评审与真实世界验证路线图。
原文摘要 · Abstract (English)
Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulness-intervention routing. Anian is not intended to diagnose psychiatric conditions or replace clinical care or crisis intervention. Its modular pipeline places generative AI downstream of structured state representation, conservative risk fusion, and response gating. User text or voice-derived ASR transcripts are mapped into four linked layers: L1 emotion states, L2 psychosocial constructs, L3 safety risk, and L4 intervention routes. Local text- and rule-based safety evidence is fused with external voice-derived evidence using a highest-risk-priority rule, S_fusion = max(S_local, S_external). At moderate or high fused risk, ordinary AI-generated responses and text-to-speech delivery are blocked and replaced by fixed safety content and prompts for human support. An internal prototype evaluation used approximately 858,295 normalized records from public emotion, dialogue, mental-health-related, and Chinese dialogue corpora within a weak-label and rule-derived framework. Micro-F1 scores were 0.9604 for L1 emotion classification, 0.9144 for L2 psychosocial constructs, and 0.9742 for L4 routing. In a controlled safety stress test of 233 samples, the L3 rule engine achieved high-risk recall of 1.0000 within predefined scenarios. These findings support the internal feasibility of the label framework and gating logic but do not establish clinical validity, diagnostic accuracy, real-world safety, or effectiveness. We report the architecture, ontology, safety-fusion mechanism, prototype evaluation, error-analysis plan, and roadmap for expert-reviewed and real-world validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。