对比12种提示设计,发现清晰定义的提示更利于检测文本极化。
Lingo_Research_Group at SemEval-2026 Task 9: Evaluating Prompt Variants for Polarization Detection

- 系统测试12种不同风格的提示,评估术语清晰度与推理引导的影响。
- 在跨语言测试中,粗粒度极化检测F1达0.762,细粒度识别仅0.444。
- 适合研究提示工程在多语言社会语用分类中的应用效果。
本论文提交至SemEval-2026任务9:多语言文本分类挑战——极化检测,涵盖三个子任务:(1) 二分类极化检测,(2) 极化类型分类,(3) 极化表现识别。我们采用系统性研究方法,设计十二种在术语清晰度、定义详尽程度、推理引导和上下文示例使用上不同的短提示。实验基于aya-101和Gemma3-27B模型进行,最终选用Gemma3-27B作为提交模型。在官方测试集上,系统在子任务1、2、3的平均宏F1分数分别为0.762、0.587、0.444,准确率分别为0.819、0.678、0.498,结果均基于22种语言的平均值。通过跨任务与跨语言分析,表明提示方法在粗粒度极化检测中有效,但在细粒度与多标签社会语言学分类中逐渐失效。
原文摘要 · Abstract (English)
Our submission presented in this paper is for SemEval-2026 Task 9: Multilingual Text Classification Challenge - Polarization Detection and it covers all three subtasks: (1) binary polarization detection, (2) polarization type classification and (3) polarization manifestation identification. We adopt a systematic approach of research on short designed prompts by considering twelve designed prompts that are different in terminology clarity, detail of the definition, guidance of reasoning and in-context examples use. The experiments are conducted using aya-101 and Gemma3-27B, with the latter chosen for the submission at the end of the development through performance considerations. Our system has an average macro level F1-score of 0.762 on Subtask 1, 0.587 on Subtask 2 and 0.444 on Subtask 3 with the average accuracy of 0.819, 0.678 and 0.498, respectively, on the official test set averaged among 22 languages, respectively. With cross-task and cross-lingual analysis, we demonstrate that prompt-based approaches can be used effectively to detect coarse grained polarization but encounter more and more difficulties as far as fine-grained and multi-label sociolinguistic classification is concerned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。