arXiv:2507.22956cs.LGcs.HC2025-07中稿 · publication at IEE…被引 5

通过敲击键盘行为检测韩语写作中使用大模型作弊,效果优于人工判断。

LLM-Assisted Cheating Detection in Korean Language via Keystrokes

  • 基于认知层次设计六类写作任务,采集用户敲击时序与节奏特征。
  • 在认知感知场景下时序特征表现更好,跨认知场景中节奏特征更稳定。
  • 模型识别转录和真实写作比改写更准,且显著超越人类判别能力。

本文提出一种基于敲击行为的框架,用于检测韩语写作中的大模型辅助作弊,弥补了以往研究在语言覆盖、认知背景及大模型介入粒度上的不足。实验包含69名参与者,在三种条件下完成写作任务:真实写作、改写ChatGPT内容、转录ChatGPT内容。每项任务涵盖布卢姆认知分类学定义的六个认知过程(记忆、理解、应用、分析、评价、创造)。提取可解释的时序与节奏特征,并在认知感知与非感知两种设定下评估多种分类器。结果显示,在认知感知场景下,时序特征表现更优;而在跨认知场景中,节奏特征更具泛化性。无论是模型还是人工评估者,识别真实写作与转录内容均比改写内容更容易,且模型性能显著优于人类。结果表明,敲击动态能可靠识别不同认知需求与写作策略下的大模型辅助写作,包括改写与转录。

原文摘要 · Abstract (English)

This paper presents a keystroke-based framework for detecting LLM-assisted cheating in Korean, addressing key gaps in prior research regarding language coverage, cognitive context, and the granularity of LLM involvement. Our proposed dataset includes 69 participants who completed writing tasks under three conditions: Bona fide writing, paraphrasing ChatGPT responses, and transcribing ChatGPT responses. Each task spans six cognitive processes defined in Bloom's Taxonomy (remember, understand, apply, analyze, evaluate, and create). We extract interpretable temporal and rhythmic features and evaluate multiple classifiers under both Cognition-Aware and Cognition-Unaware settings. Temporal features perform well under Cognition-Aware evaluation scenarios, while rhythmic features generalize better under cross-cognition scenarios. Moreover, detecting bona fide and transcribed responses was easier than paraphrased ones for both the proposed models and human evaluators, with the models significantly outperforming the humans. Our findings affirm that keystroke dynamics facilitate reliable detection of LLM-assisted writing across varying cognitive demands and writing strategies, including paraphrasing and transcribing LLM-generated responses.

作弊检测键盘行为大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。