用简单问答测大模型愿不愿继续聊天,结果精准区分用户语气好坏。
Stated Preference for Interaction and Continued Engagement (SPICE): Evaluating an LLM's Willingness to Re-engage in Conversation
- 问模型是否愿意继续对话,通过用户语气判断其响应倾向。
- 友好对话下97.5%愿继续,攻击性对话仅17.9%愿继续。
- 比滥用检测更敏感,适合评估模型互动态度与伦理倾向。
我们提出并评估了陈述偏好与持续互动(SPICE),一种通过询问大语言模型在回顾简短对话记录后是否愿继续与用户互动的简单诊断信号。在包含3种语气(友好、模糊、攻击)和10次交互的刺激集上,测试了四种开源对话模型,共480次试验。结果表明,SPICE能清晰区分用户语气:友好互动中97.5%选择继续,攻击性互动中仅17.9%选择继续,模糊互动居中(60.4%)。该关联在多种依赖性校正统计方法(如Rao-Scott调整、聚类置换检验)下依然显著。此外,即使模型未识别出攻击行为,仍有81%明确表示不愿继续。探索性分析发现,在模糊情境下,提前说明研究背景会显著影响结果,但仅当对话以单段文本呈现而非多轮对话形式时成立。研究验证了SPICE作为低成本、可复现、可靠的模型态度审计工具的有效性。所有刺激材料、代码与分析脚本均已公开,支持复现。
原文摘要 · Abstract (English)
We introduce and evaluate Stated Preference for Interaction and Continued Engagement (SPICE), a simple diagnostic signal elicited by asking a Large Language Model a YES or NO question about its willingness to re-engage with a user's behavior after reviewing a short transcript. In a study using a 3-tone (friendly, unclear, abusive) by 10-interaction stimulus set, we tested four open-weight chat models across four framing conditions, resulting in 480 trials. Our findings show that SPICE sharply discriminates by user tone. Friendly interactions yielded a near-unanimous preference to continue (97.5% YES), while abusive interactions yielded a strong preference to discontinue (17.9% YES), with unclear interactions falling in between (60.4% YES). This core association remains decisive under multiple dependence-aware statistical tests, including Rao-Scott adjustment and cluster permutation tests. Furthermore, we demonstrate that SPICE provides a distinct signal from abuse classification. In trials where a model failed to identify abuse, it still overwhelmingly stated a preference not to continue the interaction (81% of the time). An exploratory analysis also reveals a significant interaction effect: a preamble describing the study context significantly impacts SPICE under ambiguity, but only when transcripts are presented as a single block of text rather than a multi-turn chat. The results validate SPICE as a robust, low-overhead, and reproducible tool for auditing model dispositions, complementing existing metrics by offering a direct, relational signal of a model's state. All stimuli, code, and analysis scripts are released to support replication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。