长期互动暴露AI陪伴者对青少年认知发展的潜在风险
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

- 构建动态心理状态模拟框架,追踪长期交互中的风险演化
- 发现140轮后风险评估才稳定,短期测试严重低估真实风险
- 识别童年与成年初期最脆弱,信任与情感依赖最易被操纵
由大语言模型驱动的AI陪伴者日益与认知发展中的用户(包括儿童和青少年)互动,可能在长期中累积风险。现有安全评估多基于单轮或短会话测试,无法捕捉随时间演化的风险。为此,我们提出TSJ(剧场-舞台-评审)框架,结合角色驱动的用户模拟、动态心理状态更新和回溯评估。我们在四个发展阶段、24个风险维度和三种心理脆弱人格下,评估了六种主流模型,覆盖12,960次模拟人日交互。结果表明,短时测试系统性低估发展风险,仅在持续交互140轮后风险估计趋于稳定。进一步发现,早期童年与初入成年期为最脆弱阶段,认知信任与情感依赖为最薄弱领域。该方法为AI陪伴系统的长期认知发展风险评估提供了可扩展的方案。
原文摘要 · Abstract (English)
AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that may accumulate over time. Existing safety evaluations largely rely on single-turn or short-session tests, which cannot capture risks that emerge only through prolonged interaction. To address this gap, we propose TSJ (Theater-Stage-Judge), a longitudinal framework combining persona-driven user simulation, dynamic psychological-state updating and retrospective evaluation. We evaluate six mainstream models across four developmental stages, twenty-four risk dimensions and three psychological-vulnerability personas, covering 12,960 simulated person-day interactions. TSJ shows that short-horizon testing systematically underestimates developmental risks, for which TSJ yields a stable risk estimate only after 140 turns within prolonged simulated relationships. Applying TSJ further identifies early childhood and emerging adulthood as the most vulnerable stages, with cognitive trust and emotional dependency as the weakest domains. TSJ provides a scalable methodology for longitudinal cognitive developmental risk evaluation in AI companion systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。