arXiv:2608.08746cs.AI2026-08

用小模型对话高效获取经前症状评分,问答量减半仍保持高准确率。

Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models

论文配图:Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
图 1 · 摘自论文原文
  • 通过对话主动询问症状集群,用小模型还原原始严重程度标签。
  • 三问策略减少50%问题数,97.45%结果与原评分差一级内,中重度症状召回率达80.94%。
  • 适合需要轻量、可部署的日常健康追踪系统,尤其关注用户体验的场景。

前瞻性每日症状追踪是经前健康评估的核心,但重复的序数表单带来沉重负担。本文将对话式管理建模为序数标签恢复问题:系统主动询问少量症状集群,并将回应映射回原始严重度标签。基于mcPHASES数据集的3,320个完整参与者日,涵盖腹痛、情绪波动、疲劳、睡眠问题、压力和胀气,采用六级评分。六名参与者用于开发,36名用于冻结评估,共360个参与者日和2,160个条目标签。ModernBERT证据门判断症状是否表达,Qwen2.5-1.5B-Instruct生成确定性结构化严重度评分。固定六项提问达成0.976的加权肯德尔等级相关系数,而三项联合症状集群提问达0.913,97.45%一致在一级内,中重度症状召回率80.94%,问题数减少50%。开放优先自适应策略需3.92–5.98次提问,一致性低于对应固定策略。参与者集群自助分析估计三集群与六项策略间κ值差异为-0.062(95%置信区间-0.076至-0.048)。主动集群级问询为自然对话到可复用日度症状标签提供了直接、本地化的模型路径。

原文摘要 · Abstract (English)

Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden. We formulate conversational administration as an ordinal label-recovery problem: the system actively elicits a small set of symptom clusters and maps each response to the original severity labels. We used 3,320 complete participant-days from the mcPHASES dataset, covering cramps, mood swing, fatigue, sleep issues, stress, and bloating on a six-level scale. Six participants were reserved for development and 36 for a frozen evaluation comprising 360 participant-days and 2,160 item labels. A ModernBERT evidence gate detected whether a symptom was expressed, and Qwen2.5-1.5B-Instruct produced deterministic structured severity scores. Fixed six-item questioning achieved a quadratic weighted kappa of 0.976, whereas three joint symptom-cluster questions achieved 0.913, 97.45% agreement within one severity level, and 80.94% recall for moderate-or-higher symptoms while reducing questions by 50%. Open-first adaptive policies required 3.92-5.98 questions and produced lower agreement than the corresponding fixed policies. Participant-cluster bootstrap analysis estimated a kappa difference of -0.062 (95% CI -0.076 to -0.048) between the three-cluster and six-item strategies. Active cluster-level elicitation provides a direct, local-model route from natural conversation to reusable daily symptom labels.

健康追踪小模型对话系统症状评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。