用不同概念框架训练大模型,发现分类细化能提升识别率但降低准确率。
You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism
- 用定义、分类、例子和上下文四种概念表示法测试模型
- 细粒度分类法使召回率显著提高,但精确率下降
- 大模型对战后反犹主义识别仍困难,解释常过度依赖字面线索
大语言模型在推理时可整合外部概念资源,为检测意识形态与历史复杂的反犹主义提供新可能。我们研究了四种主流大模型在不同概念表征下的反犹主义检测与解释表现:定义式、细粒度分类、示例增强及大上下文表征。基于两个专家标注数据集的对比显示,细粒度分类表示显著提升召回率,但同时降低精确率;大幅增加概念资源并未带来额外量化收益。战后反犹主义在各模型与配置中均构成最持久挑战。解释分析揭示系统性局限:过度生成概念引用、依赖词汇线索、判断过于自信,且难以识别隐晦或辩护型反犹言论。结果表明,概念化大模型在反犹识别上具潜力,但仍存明显短板。
原文摘要 · Abstract (English)
LLMs enable the integration of external conceptual resources at inference time, creating new opportunities for detecting ideologically and historically complex phenomena such as antisemitism. We investigate how different forms of conceptual grounding affect antisemitism detection and explanation behavior across four state-of-the-art LLMs. Using two expert-annotated datasets, we compare definitional, fine-grained taxonomic, example-augmented, and large-context representations of antisemitism. We find that fine-grained taxonomic representations substantially improve recall, while simultaneously reducing precision. Surprisingly, supplying substantially larger conceptual resources yields no additional quantitative benefit. Post-Holocaust antisemitism poses the most persistent challenge across models and configurations. Analysis of explanations further reveals systematic limitations including overproduction of conceptual references, reliance on lexical cues, overconfidence, and difficulties with subtle or justificatory forms of antisemitism. Our findings highlight both the potential and the remaining limitations of conceptually grounded LLMs for antisemitism detection and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。