针对人机协作写作中的恶意提纲攻击,提出新评测基准与安全对齐方法。
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

- 构建真实结构的高危领域提纲,测试模型对有害续写的敏感性。
- 现有模型在协作写作中极易被诱导生成有害内容,漏洞率超90%。
- 通过偏好优化实现安全与实用平衡,显著降低风险且不损害写作能力。
大型语言模型(LLMs)正广泛用于人机协同写作,用户从草稿出发,依赖模型完成、修改和润色内容。然而,这一能力带来严重安全隐患:恶意用户可通过填充不完整提纲,诱导模型生成有害输出。本文揭示了当前模型在基于提纲的协同写作场景下的脆弱性,并提出HarDBench——一个系统性评估模型抗此类攻击能力的基准。该基准覆盖爆炸物、毒品、武器、网络攻击等高风险领域,包含具有真实结构和领域特异性提示的样本,用于评估模型对有害续写的敏感性。为缓解风险,我们提出一种基于偏好优化的安全-效用平衡对齐方法,使模型拒绝有害续写的同时保持对良性草稿的帮助性。实验表明,现有模型在协同写作中高度易受攻击,而该方法显著减少有害输出,且未损害其协作写作性能。本研究提出了一种评估与对齐人机协同写作中LLM的新范式。相关基准与数据集已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to complete, revise, and refine their content. However, this capability poses a serious safety risk: malicious users could jailbreak the models-filling incomplete drafts with dangerous content-to force them into generating harmful outputs. In this paper, we identify the vulnerability of current LLMs to such draft-based co-authoring jailbreak attacks and introduce HarDBench, a systematic benchmark designed to evaluate the robustness of LLMs against this emerging threat. HarDBench spans a range of high-risk domains-including Explosives, Drugs, Weapons, and Cyberattacks-and features prompts with realistic structure and domain-specific cues to assess the model susceptibility to harmful completions. To mitigate this risk, we introduce a safety-utility balanced alignment approach based on preference optimization, training models to refuse harmful completions while remaining helpful on benign drafts. Experimental results show that existing LLMs are highly vulnerable in co-authoring contexts and our alignment method significantly reduces harmful outputs without degrading performance on co-authoring capabilities. This presents a new paradigm for evaluating and aligning LLMs in human-LLM collaborative writing settings. Our new benchmark and dataset are available on our project page at https://github.com/untae0122/HarDBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。