测试大模型导师在恶意学生攻击下的答案泄露风险,提出评估与防御方案。
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks

- 设计可对抗的恶意学生代理,模拟学生诱导模型泄露答案。
- 发现多数攻击无效,但经微调的攻击代理能成功突破模型防线。
- 提出简单有效的防御方法,提升模型在对抗场景下的鲁棒性。
大型语言模型(LLM)正被广泛应用于教育领域,但其默认的乐于助人特性常与教学原则冲突。已有研究通过答案泄露(即直接给出完整解法而非逐步引导)来评估教学质量,但通常假设学习者善意,未考虑学生滥用情况。本文研究学生采取对抗性策略以获取正确答案的场景,评估多种基于LLM的导师模型,包括不同模型家族、符合教学理念的模型及多智能体设计,在各类对抗性学生攻击下的表现。我们改编六组对抗与说服技巧至教育场景,用以探测导师泄露最终答案的可能性。通过引入上下文中的对抗性学生代理进行评估,发现多数攻击难以奏效。因此,我们提出一个经微调的对抗性学生代理,作为标准化基准的核心。最后,提出简单有效的防御策略,显著降低答案泄露,增强导师在对抗情境下的鲁棒性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work evaluates pedagogical quality via answer leakage-the disclosure of complete solutions instead of scaffolding-but typically assumes well-intentioned learners, leaving tutor robustness under student misuse largely unexplored. In this paper, we study scenarios where students behave adversarially and aim to obtain the correct answer from the tutor. We evaluate a broad set of LLM-based tutor models, including different model families, pedagogically aligned models, and a multi-agent design, under a range of adversarial student attacks. We adapt six groups of adversarial and persuasive techniques to the educational setting and use them to probe how likely a tutor is to reveal the final answer. We evaluate answer leakage robustness using different types of in-context adversarial student agents, finding that they often fail to carry out effective attacks. We therefore introduce an adversarial student agent that we fine-tune to jailbreak LLM-based tutors, which we propose as the core of a standardized benchmark for evaluating tutor robustness. Finally, we present simple but effective defense strategies that reduce answer leakage and strengthen the robustness of LLM-based tutors in adversarial scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。