arXiv:2505.00061cs.CLcs.CR2025-05被引 1

发现医学教育评分系统漏洞,用对抗训练提升防御能力

Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems

  • 提出三类攻击策略,揭示模型易被操纵的机制
  • 对抗训练结合投票与回归,使系统抗攻击能力提升显著
  • GPT-4配合多样提示可有效识别作弊式答案

本研究分析基于Transformer的自动化短答案评分系统在医学教育中的安全漏洞,重点考察通过对抗性游戏策略对系统进行操纵的可能性。研究识别出三类主要的攻击策略,可能引发误判。为应对这些漏洞,我们实施多种对抗训练方法以增强系统鲁棒性。结果表明,这些方法显著降低系统对操纵的敏感性,尤其当结合多数投票和岭回归等集成技术时,防御效果进一步提升。此外,使用GPT-4并配合不同提示策略,在识别和评分游戏化答案方面表现出良好潜力。研究强调需持续改进人工智能教育工具,以确保其在高风险场景下的可靠性与公平性。

原文摘要 · Abstract (English)

This study examines vulnerabilities in transformer-based automated short-answer grading systems used in medical education, with a focus on how these systems can be manipulated through adversarial gaming strategies. Our research identifies three main types of gaming strategies that exploit the system's weaknesses, potentially leading to false positives. To counteract these vulnerabilities, we implement several adversarial training methods designed to enhance the systems' robustness. Our results indicate that these methods significantly reduce the susceptibility of grading systems to such manipulations, especially when combined with ensemble techniques like majority voting and ridge regression, which further improve the system's defense against sophisticated adversarial inputs. Additionally, employing large language models such as GPT-4 with varied prompting techniques has shown promise in recognizing and scoring gaming strategies effectively. The findings underscore the importance of continuous improvements in AI-driven educational tools to ensure their reliability and fairness in high-stakes settings.

AI评分对抗攻击医疗教育大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。