用智能体框架自动评估心理健康筛查中的提示行为,提升稳定性与可审计性。
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening

- 构建多智能体系统,自动设计、评估和选择提示策略。
- 在访谈文本中实现抑郁症筛查的稳定表现,保障行为一致性。
- 适合关注大模型在医疗对话中安全可控的开发者与研究者。
本文提出CPEMH,一种面向基于转录文本的心理健康筛查基础模型系统的提示驱动行为评估与保证的智能体框架。该框架作为大规模语言系统的行为保障工程方法,引入协同架构,自主完成提示策略的设计、评估与选择,实现跨场景行为变异性系统化控制。其模块化智能体设计融合编排器、推理与评估智能体,确保提示生命周期中的可追溯性、可复现性与鲁棒性。在自动抑郁筛查的案例研究中,验证了框架在对话与临床敏感领域稳定并审计基础模型行为的能力。经验表明:模块化编排对行为保障至关重要;稳定性应优先于架构复杂度;F1值、偏差与鲁棒性应作为核心验收标准。
原文摘要 · Abstract (English)
This paper presents CPEMH, an agentic framework designed to evaluate prompt-driven behavior in foundation-model systems operating on transcript-based datasets for mental-health screening. CPEMH serves as an engineering methodology for behavioral assurance in large-scale language systems, introducing an orchestrated architecture that autonomously performs the design, evaluation, and selection of prompt strategies, enabling systematic control of behavioral variability across contexts. Its modular agentic design, combining orchestrator, inference, and evaluation agents, ensures traceability, reproducibility, and robustness throughout the prompting lifecycle. A case study on automated depression screening from interview transcripts demonstrates the framework's capacity to stabilize and audit foundation-model behavior in conversational and clinically sensitive domains. Lessons learned emphasize the role of modular orchestration in behavioral assurance, the prioritization of stability over architectural complexity, and the integration of F1, bias, and robustness as core acceptance criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。