测试零样本思维链对日语提示的效果,发现高级模型表现反而下降。
Effectiveness of Zero-shot-CoT in Japanese Prompts
- 在日语提示中加入'一步步思考'引导推理
- GPT-3.5部分任务提升,GPT-4o-mini整体性能下降
- 大学数学和抽象代数仍有效,适合日语推理研究者
我们使用ChatGPT-3.5和4o-mini,对比了零样本思维链(zero-shot CoT)提示在日语与英语中的效果。该技术通过在提示末尾添加如'让我们一步一步思考'的短语,以促进模型推理,已在英文数学与推理任务中证明有效。本文采用日本多任务语言理解基准(JMMLU)和多任务语言理解基准(MMLU)评估其在日语中的迁移效果。结果显示,尽管零样本CoT在GPT-3.5的部分提示类别中带来显著性能提升,但在GPT-4o-mini中却导致整体性能明显下降。然而,在日语提示中,大学数学和抽象代数等特定类别仍保持改进,表明其在某些任务上仍有价值。
原文摘要 · Abstract (English)
We compare the effectiveness of zero-shot Chain-of-Thought (CoT) prompting in Japanese and English using ChatGPT-3.5 and 4o-mini. The technique of zero-shot CoT, which involves appending a phrase such as "Let's think step by step" to a prompt to encourage reasoning before answering, has been shown to offer LLM performance improvements in mathematical and reasoning tasks, particularly in English. We investigate how these effects transfer to Japanese using the Japanese Multi-task Language Understanding Benchmark (JMMLU) and the Multi-task Language Understanding Benchmark (MMLU). Our results show that while zero-shot CoT prompting can lead to notable performance gains for some prompt categories in GPT-3.5, its impact in GPT-4o-mini is associated with significant performance declines. However, for Japanese prompts there remain certain categories, such as college mathematics and abstract algebra, that still exhibit improvements, despite the broader trend of diminishing effectiveness in more advanced models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。