让AI文本读起来像真人写的,用小模型效果更好。
Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer
- 用三个模型对比改写效果,小模型反而更接近人类风格。
- 最大模型虽然风格变化大,但不准确,易过度修改。
- 提出新评估标准:风格变化准确性比幅度更重要。
AI生成文本在学术和专业写作中已很常见,促使研究者探索检测方法。而反向问题——将AI文本系统性重写为自然的人类写作风格——仍较少被研究。本文构建了一个包含25,140对的平行语料库,识别出11个可量化的风格标记来区分AI与人类文本,并微调了BART-base、BART-large和Mistral-7B-Instruct(QLoRA)三个模型。实验显示,参数量仅为Mistral-7B的1/17,BART-large在参考文本相似度上表现最优:BERTScore F1达0.924,ROUGE-L为0.566,chrF++为55.92。尽管Mistral-7B在风格标记迁移上得分更高,但其表现实为过度修正而非真实准确。研究指出,当前风格转换评估中对‘迁移准确性’的关注不足,应引入更合理的评价维度。
原文摘要 · Abstract (English)
AI-generated text has become common in academic and professional writing, prompting research into detection methods. Less studied is the reverse: systematically rewriting AI-generated prose to read as genuinely human-authored. We build a parallel corpus of 25,140 paired AI-input and human-reference text chunks, identify 11 measurable stylistic markers separating the two registers, and fine-tune three models: BART-base, BART-large, and Mistral-7B-Instruct with QLoRA. BART-large achieves the highest reference similarity -- BERTScore F1 of 0.924, ROUGE-L of 0.566, and chrF++ of 55.92 -- with 17x fewer parameters than Mistral-7B. We show that Mistral-7B's higher marker shift score reflects overshoot rather than accuracy, and argue that shift accuracy is a meaningful blind spot in current style transfer evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。