构建1万条精神分裂症风险症状合成访谈数据,支持证据溯源的临床评估研究。
AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

- 基于真实量表生成结构化合成访谈,每轮对话均有文本依据。
- 10,000条数据覆盖24个症状项与最终衰减精神病综合征诊断。
- 专为评估模型提取证据、追溯支持语句能力设计,适合临床AI研究者。
精神分裂症风险评估的进展受限于数据获取瓶颈。真实临床访谈因隐私、治理和知情同意问题难以共享。本文提出AnchorSIPS,一个包含10,000条结构化精神病风险访谈的合成数据集,其内容基于Mini-SIPS量表。每条访谈涵盖病史、24个症状问题、患者确认项的后续证据、妄想类症状(异常信念)、幻觉类症状(异常感知)及紊乱沟通的判断,排除明确精神病症状(“明显精神病”),并最终判定衰减精神病综合征(APS)状态——一种轻度或早期精神病症状的高风险状态。该诊断非独立标签,而是依赖前述确认项、支持性细节、症状分类决策与精神病排除检查。所有中间决策均锚定至原始对话片段。数据通过规划-实现流水线生成:隐藏病例表定义患者临床状态,确定性规划器固定访谈结构,大模型仅生成受验证与边界修复的患者话语。先固定标签与结构避免多轮对话中常见的前后不一致。在七种大模型基线测试中,模型能恢复粗粒度判断,但无法提取后续细节或引用支持语句,导致最终标签性能虚高。AnchorSIPS旨在推动证据抽取、基于转录文本的测量与部分披露下的不确定性研究。
原文摘要 · Abstract (English)
Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。