微调提示词竟让AI情感分类结果大变,可信度存疑
Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification
- 用10组微调提示词测试GPT-4o mini对万条西语评论的情感分类
- 多数提示差异导致分类结果显著不同,个别情况甚至混类或乱答
- 提示不结构化会大幅增加幻觉,警示研究者谨慎使用大模型
当前社会科学面临一个核心问题:我们能在多大程度上信任像ChatGPT这样复杂的预测模型?本研究检验了假设——微小的提示结构调整不会对GPT-4o mini在情感极性分析任务中的分类结果产生显著影响。基于包含10万条关于四位拉丁美洲总统的西语文本数据集,模型在10次实验中对每条评论进行正/负/中性分类,每次仅微调提示词。通过探索性和验证性分析,发现即使词汇、句法或语气上的细微变化,也会显著影响分类结果。部分情况下模型出现类别混淆、附加解释或使用非西班牙语回应。卡方检验显示,多数提示对比间存在显著差异,仅当语言结构高度相似时无显著变化。研究挑战了大模型在分类任务中的鲁棒性,揭示其对指令变化的高度敏感性。此外,提示缺乏结构化会增加幻觉频率。讨论指出,对大模型的信任不仅取决于技术性能,更依赖于其使用背后的社会与制度关系。
原文摘要 · Abstract (English)
One fundamental question for the social sciences today is: how much can we trust highly complex predictive models like ChatGPT? This study tests the hypothesis that subtle changes in the structure of prompts do not produce significant variations in the classification results of sentiment polarity analysis generated by the Large Language Model GPT-4o mini. Using a dataset of 100.000 comments in Spanish on four Latin American presidents, the model classified the comments as positive, negative, or neutral on 10 occasions, varying the prompts slightly each time. The experimental methodology included exploratory and confirmatory analyses to identify significant discrepancies among classifications. The results reveal that even minor modifications to prompts such as lexical, syntactic, or modal changes, or even their lack of structure impact the classifications. In certain cases, the model produced inconsistent responses, such as mixing categories, providing unsolicited explanations, or using languages other than Spanish. Statistical analysis using Chi-square tests confirmed significant differences in most comparisons between prompts, except in one case where linguistic structures were highly similar. These findings challenge the robustness and trust of Large Language Models for classification tasks, highlighting their vulnerability to variations in instructions. Moreover, it was evident that the lack of structured grammar in prompts increases the frequency of hallucinations. The discussion underscores that trust in Large Language Models is based not only on technical performance but also on the social and institutional relationships underpinning their use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。