LLM在零样本信念预测任务上表现不佳,新方法显著提升效果
Zero-Shot Belief: A Hard Problem for LLMs
- 统一框架+微调DeBERTa联合检测事件与信念标签
- 在FactBank上达到新最好结果,但多数LLM仍表现差
- 验证了方法在意大利语数据集ModaFact上的泛化能力
我们提出两种基于大模型的零样本源与目标信念预测方法:一种是单次完成事件、来源和信念标签识别的统一系统;另一种是结合微调DeBERTa分词器进行事件检测的混合方法。实验表明,多种开源、闭源及基于推理的LLM在此任务上均表现不佳。采用混合方法后,我们在FactBank数据集上取得新的最佳性能,并进行了详尽的错误分析。该方法进一步在意大利语信念语料库ModaFact上进行了测试,验证了其跨语言泛化能力。
原文摘要 · Abstract (English)
We present two LLM-based approaches to zero-shot source-and-target belief prediction on FactBank: a unified system that identifies events, sources, and belief labels in a single pass, and a hybrid approach that uses a fine-tuned DeBERTa tagger for event detection. We show that multiple open-sourced, closed-source, and reasoning-based LLMs struggle with the task. Using the hybrid approach, we achieve new state-of-the-art results on FactBank and offer a detailed error analysis. Our approach is then tested on the Italian belief corpus ModaFact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。