研究政治文本中价值观检测的上下文与道德知识作用,发现大模型不必然更优。
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
- 对比句子、文档级上下文和检索增强,验证信息范围影响
- 完整文档提升小模型效果3.8-4.8点,但对大模型帮助有限
- 引入道德知识库可稳定增益,尤其对易混淆价值观有效
在政治文本中检测施瓦茨价值观面临挑战,因隐含线索依赖上下文及细微价值差异。本研究系统评估上下文长度、显式道德知识对句级价值识别的影响。采用ValuesML/Touché ValueEval标准,比较句子、窗口与全文输入;无RAG与带自建道德知识库的检索增强设置;监督式DeBERTa-v3-base/large编码器;以及12B至123B参数的零样本大语言模型。结果表明:更多上下文并非总是更好——全文档输入使监督型DeBERTa模型宏F1提升3.8–4.8点,但对零样本大模型无一致增益。检索到的道德知识在匹配对比中更具稳定性,提升所有测试模型家族和上下文条件下早期融合的表现。模型规模从DeBERTa-v3-base扩展至large,或从12B扩展至更大参数量,并未保证性能提升。简单早期融合优于测试中的晚期融合与交叉注意力RAG变体。逐值分析显示,上下文与检索对社会嵌入性或概念上易混淆的价值最有益。研究建议:价值敏感的NLP应联合评估上下文、知识与模型族,而非默认更长输入或更大模型即为最优。
原文摘要 · Abstract (English)
Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions between neighboring values. We study when context and explicit moral knowledge help sentence-level value detection. Using the ValuesML/Touché ValueEval format, we compare sentence, window, and full-document inputs; no-RAG and retrieval-augmented settings with a curated moral knowledge base; supervised DeBERTa-v3-base/large encoders; and zero-shot LLMs from 12B to 123B parameters. The results show that more context is not uniformly better: full-document context improves supervised DeBERTa encoders by 3.8-4.8 macro-F1 points over sentence-only input, but does not consistently help zero-shot LLMs. Retrieved moral knowledge is more consistently useful in matched comparisons, improving each tested model family and context condition under early fusion. However, scaling from DeBERTa-v3-base to large and from 12B to larger LLMs does not guarantee gains, and simple early fusion outperforms the tested late-fusion and cross-attention RAG variants for encoders. Per-value analyses show that context and retrieval help most for socially situated or conceptually confusable values. These findings suggest that value-sensitive NLP should evaluate context, knowledge, and model family jointly rather than treating longer inputs or larger models as universal improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。