arXiv:2601.08176cs.CLcs.AI2026-01被引 3

通过提示工程提升大模型在政治问答中的清晰度与主题识别能力

Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering

  • 采用思维链+少样本提示,显著提升清晰度判断准确率
  • 清晰度预测准确率达63%,主题识别准确率从60%升至74%
  • 结构化提示对高阶评估有效,细粒度逃避行为仍难捕捉

大语言模型在政治问答中的自动评估不仅需事实正确,还需保证表达清晰。本文基于SemEval 2026共享任务的CLARITY数据集,研究提示设计对清晰度评估的影响。对比GPT-3.5基线与GPT-5.2在三种提示策略下的表现:简单提示、思维链提示、思维链+少样本提示。结果表明,使用思维链+少样本提示时,清晰度预测准确率由56%提升至63%;思维链提示在逃避行为识别上达到最高准确率34%,但细粒度类别间表现不稳定。主题识别方面,基于推理的提示使准确率从60%提升至74%。整体表明提示设计可有效提升高层级清晰度评估,但细粒度逃避检测与主题识别仍具挑战。

原文摘要 · Abstract (English)

Automatic evaluation of large language model (LLM) responses requires not only factual correctness but also clarity, particularly in political question-answering. While recent datasets provide human annotations for clarity and evasion, the impact of prompt design on automatic clarity evaluation remains underexplored. In this paper, we study prompt-based clarity evaluation using the CLARITY dataset from the SemEval 2026 shared task. We compare a GPT-3.5 baseline provided with the dataset against GPT-5.2 evaluated under three prompting strategies: simple prompting, chain-of-thought prompting, and chain-of-thought with few-shot examples. Model predictions are evaluated against human annotations using accuracy and class-wise metrics for clarity and evasion, along with hierarchical exact match. Results show that GPT-5.2 consistently outperforms the GPT-3.5 baseline on clarity prediction, with accuracy improving from 56 percent to 63 percent under chain-of-thought with few-shot prompting. Chain-of-thought prompting yields the highest evasion accuracy at 34 percent, though improvements are less stable across fine-grained evasion categories. We further evaluate topic identification and find that reasoning-based prompting improves accuracy from 60 percent to 74 percent relative to human annotations. Overall, our findings indicate that prompt design reliably improves high-level clarity evaluation, while fine-grained evasion and topic detection remain challenging despite structured reasoning prompts.

大模型评估提示工程政治问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。