让对话模型在助人时避免控制用户,保持用户自主性。
Care-Conditioned Neuromodulation for Autonomy-Preserving Supportive Dialogue Agents
- 用用户状态动态调节回复生成,防止过度干预
- 在多轮对话中提升自主性支持得分0.25分
- 适合情感支持、心理咨询等敏感场景
部署于支持性或咨询性角色的大语言模型需在帮助用户与保护其自主性之间取得平衡,但现有对齐方法主要优化帮助性和无害性,未显式建模依赖强化、过度保护或强迫引导等关系风险。本文提出关怀条件神经调制(CCN),一种基于状态的控制框架:从用户状态和对话上下文中学习标量信号,用于调节响应生成与候选选择。将该问题形式化为自主性保护对齐任务,定义效用函数,奖励自主性支持与帮助性,惩罚依赖与强制。构建包含反复安慰依赖、操纵性关怀、过度保护和边界不一致的多轮对话关系失效模式基准。在该基准上,结合关怀条件候选生成与效用重排序,在自主性保护效用上优于监督微调+0.25,优于偏好优化+0.07,同时保持相近的支持性水平。初步人类评估及零样本迁移至真实情绪支持对话显示结果与自动指标方向一致。结果表明,状态依赖控制结合效用选择是实现敏感对话多目标对齐的可行路径。
原文摘要 · Abstract (English)
Large language models deployed in supportive or advisory roles must balance helpfulness with preservation of user autonomy, yet standard alignment methods primarily optimize for helpfulness and harmlessness without explicitly modeling relational risks such as dependency reinforcement, overprotection, or coercive guidance. We introduce Care-Conditioned Neuromodulation (CCN), a state-dependent control framework in which a learned scalar signal derived from structured user state and dialogue context conditions response generation and candidate selection. We formalize this setting as an autonomy-preserving alignment problem and define a utility function that rewards autonomy support and helpfulness while penalizing dependency and coercion. We also construct a benchmark of relational failure modes in multi-turn dialogue, including reassurance dependence, manipulative care, overprotection, and boundary inconsistency. On this benchmark, care-conditioned candidate generation combined with utility-based reranking improves autonomy-preserving utility by +0.25 over supervised fine-tuning and +0.07 over preference optimization baselines while maintaining comparable supportiveness. Pilot human evaluation and zero-shot transfer to real emotional-support conversations show directional agreement with automated metrics. These results suggest that state-dependent control combined with utility-based selection is a practical approach to multi-objective alignment in autonomy-sensitive dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。