arXiv:2511.16445cs.AIcs.LG2025-11ICML被引 1

构建聊天记录异常检测基准,追踪痴呆患者语言行为随时间的渐变。

PersonaDrift: A Benchmark for Temporal Anomaly Detection in Language-Based Dementia Monitoring

  • 基于照护者访谈生成60天模拟对话数据,模拟语言表达减弱与离题变化。
  • 情感平淡化可用简单统计模型检测,语义偏移需时序建模和个性化基线。
  • 个性化模型显著优于通用模型,凸显个体行为背景的重要性。

患有痴呆症的人群(PLwD)在沟通中常出现渐进性变化,如表达减少、重复增多或话题偏离等微妙特征。尽管照护者可能察觉这些变化,但多数计算工具无法长期追踪此类行为漂移。本文提出PersonaDrift,一个基于真实照护者访谈构建的合成基准,用于评估机器学习与统计方法在检测日常交流中渐进性改变的能力。该基准模拟60天内针对数字提醒系统的交互日志,涵盖60名基于真实案例生成的虚拟用户,其性格、语气和沟通习惯各异。重点聚焦照护者强调的两类纵向变化:情感平淡化(情绪强度与表达量下降)和离题回复(语义漂移)。这两类变化以不同速率逐步注入,模拟自然认知发展轨迹。框架设计支持未来扩展至其他行为模式。我们评估了多种异常检测方法:无监督统计方法(CUSUM、EWMA、One-Class SVM)、使用上下文嵌入的序列模型(GRU + BERT)以及广义与个性化设置下的监督分类器。初步结果表明,在基线变化小的用户中,情感平淡化可被简单统计模型有效捕捉;而语义漂移则需依赖时序建模与个性化基准。在两项任务中,个性化分类器均显著优于通用模型,凸显个体行为上下文的关键作用。

原文摘要 · Abstract (English)

People living with dementia (PLwD) often show gradual shifts in how they communicate, becoming less expressive, more repetitive, or drifting off-topic in subtle ways. While caregivers may notice these changes informally, most computational tools are not designed to track such behavioral drift over time. This paper introduces PersonaDrift, a synthetic benchmark designed to evaluate machine learning and statistical methods for detecting progressive changes in daily communication, focusing on user responses to a digital reminder system. PersonaDrift simulates 60-day interaction logs for synthetic users modeled after real PLwD, based on interviews with caregivers. These caregiver-informed personas vary in tone, modality, and communication habits, enabling realistic diversity in behavior. The benchmark focuses on two forms of longitudinal change that caregivers highlighted as particularly salient: flattened sentiment (reduced emotional tone and verbosity) and off-topic replies (semantic drift). These changes are injected progressively at different rates to emulate naturalistic cognitive trajectories, and the framework is designed to be extensible to additional behaviors in future use cases. To explore this novel application space, we evaluate several anomaly detection approaches, unsupervised statistical methods (CUSUM, EWMA, One-Class SVM), sequence models using contextual embeddings (GRU + BERT), and supervised classifiers in both generalized and personalized settings. Preliminary results show that flattened sentiment can often be detected with simple statistical models in users with low baseline variability, while detecting semantic drift requires temporal modeling and personalized baselines. Across both tasks, personalized classifiers consistently outperform generalized ones, highlighting the importance of individual behavioral context.

痴呆监测语言分析异常检测时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。