对比人类与大模型文本风格差异,建立可解释的评测基准。
Benchmark of stylistic variation in LLM-generated texts
- 用多维度分析法比较人类与大模型文本风格
- 发现指令微调模型更接近人类语言模式
- 提供跨语言、可量化的模型风格评测标准
本研究探究人类写作与大语言模型生成文本在语体风格上的差异。采用比伯多维分析(MDA)方法,对一组人类写作文本及其对应的AI生成文本进行分析,识别出大模型与人类在哪些语言维度上存在显著且系统性的差异。使用新构建的AI-Brown语料库(与当代英式英语的BE-21语料库可比),并针对捷克语重复相同分析,使用AI-Koditex语料库和捷克多维模型。共评估16个前沿大模型在不同设置与提示下的表现,重点关注基础模型与指令微调模型的差异。基于此,建立了一个可解释的评测基准,用于模型间的横向比较与排序。
原文摘要 · Abstract (English)
This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-created texts generated to be their counterparts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. As textual material, a new LLM-generated corpus AI-Brown is used, which is comparable to BE-21 (a Brown family corpus representing contemporary British English). Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech using AI-Koditex corpus and Czech multidimensional model. Examined were 16 frontier models in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. Based on this, a benchmark is created through which models can be compared with each other and ranked in interpretable dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。