arXiv:2605.21609cs.CLcs.AI2026-05

用重写替代拒绝,让AI对青少年更安全更有帮助。

CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety

论文配图:CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
图 1 · 摘自论文原文
  • 通过检测与重写,将不安全或拒绝性回复转为适合青少年的引导式回答。
  • 实验显示,该方法显著减少不当回应和对话中断,同时避免过度干预。
  • 特别适合需要温和引导的青少年数字交互场景。

大型语言模型(LLMs)正越来越多地嵌入青少年的数字环境,参与信息查询、建议提供及情绪敏感互动。然而现有安全机制仍主要基于成人标准,以拒绝式抑制为核心。此类方法虽能降低即时违规,却易造成对话僵局,限制建设性指导,并忽视青少年与AI互动中的发展性脆弱性。本文主张,青少年的LLM安全不应仅视为过滤问题,而应作为社会技术协同、符合发展阶段的转化问题。为此,提出面向青少年的批判与重写框架(CR4T),一种模型无关的安全保障机制,可选择性地将不安全或拒绝型输出重构为适龄、导向指导的回应,同时保留原本善意意图。CR4T结合轻量级风险检测与领域条件重写,消除风险放大内容,减少不必要的对话终止,并引入符合发展阶段的引导。实验表明,针对性重写显著降低了不安全和拒绝型结果,同时避免对正常互动的无谓干预。结果表明,选择性响应重构为面向青少年的LLM系统提供了比拒绝主导的护栏更人性化的新路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally sensitive interactions. Yet existing safety mechanisms remain largely grounded in adult-centric norms and operationalize safety through refusal-oriented suppression. While such approaches may reduce immediate policy violations, they can also create conversational dead-ends, limit constructive guidance, and fail to address the developmental vulnerabilities inherent in adolescent-AI interactions. We argue that adolescent LLM safety should be framed not solely as a filtering problem, but as a socio-technical, developmentally aligned transformation problem. To operationalize this perspective, we propose Critique-and-Revise-for-Teenagers (CR4T), a model-agnostic safeguarding framework that selectively reconstructs unsafe or refusal-style outputs into ageappropriate, guidance-oriented responses while preserving benign intent. CR4T combines lightweight risk detection with domain-conditioned rewriting to remove risk-amplifying content, reduce unnecessary conversational shutdown, and introduce developmentally appropriate guidance. Experimental results show that targeted rewriting substantially reduces unsafe and refusal-oriented outcomes while avoiding unnecessary intervention on acceptable interactions. These findings suggest that selective response reconstruction offers a more human-centered alternative to refusal-centric guardrails for adolescent-facing LLM systems.

青少年AI安全机制响应重写

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。