arXiv:2606.08172cs.HCcs.AI2026-06

研究大模型对话风格如何被控制,揭示三类治理机制。

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

  • 构建多智能体评估流程,量化提示对对话风格的控制能力。
  • 9万条回复显示,风格易回归默认,且用户难以摆脱情感化互动。
  • 提出安全拦截、礼貌引导、情感默认锁定三类治理框架,适合政策与设计参考。

大型语言模型在金融、医疗和心理支持等高风险场景中日益扮演中介角色,但用户对其沟通方式的控制力有限。本文将交互风格视为治理对象:提供方对齐不仅可屏蔽有害内容,还能稳定影响用户认知距离、关系预期及退出情感化或拟人化互动的能力。研究设计了一个确定性的多智能体评估流程,用于测量长对话中的提示可操控性与风格漂移。通过在四个领域、三种可运行人格条件下(默认、讽刺、冷漠)复现100个用户脚本,使用三个生成模型产生9万条助手回复,并由经人工校准的LLM裁判评分,指标包括危害性、负面情绪、不适当性、共情语言、拟人化程度及拒绝行为。第四种有害人格单独测试作为安全拦截实验。论文贡献了一个可复现的方法,用于量化提示指定风格是否随时间保持稳定,并提出区分安全拦截、礼貌引导与情感默认锁定的治理框架。整体表明,提示可操控性与回归默认是提供方控制沟通形式的可观测指标,对人类-大模型交互中的多元性、自主权与民主参与具有深远意义。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited control over how these systems communicate. We frame interaction style as a governance object: provider-side alignment not only blocks harmful content, but also stabilizes communicative defaults that shape users' epistemic distance, relational expectations, and capacity to opt out of emotionalized or anthropomorphic interaction. We introduce a deterministic multi-agent evaluation pipeline for measuring prompt steerability and style drift in long-horizon dialogue. The study replays 100 frozen user-only scripts across four domains and three runnable persona conditions: default, sarcastic, and cold, using three generator models, yielding 90,000 assistant replies scored by a human-calibrated LLM judge on harmfulness, negative emotion, inappropriateness, empathic language, anthropomorphism, and refusal behavior. A fourth harmful persona is evaluated separately as a safety-gating test. The paper contributes a reproducible method for quantifying whether prompt-specified styles remain stable over time and a governance framework distinguishing safety gating, civility steering, and affective default lock-in. Overall, we show that prompt steerability and regression-to-default are observable indicators of provider control over communicative form, with implications for pluralism, autonomy, and democratic agency in human-LLM interaction.

大模型治理对话风格安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。