用简单规则选上下文,小模型也能高效分类叙事角色。
LTG at SemEval-2025 Task 10: Optimizing Context for Classification of Narrative Roles
- 基于实体导向的启发式方法选取关键上下文片段。
- 在有限上下文窗口下达到或超越大模型微调效果。
- 适合资源受限但需精准角色分类的场景。
我们在 SemEval 2025 共享任务 10 的子任务 1(实体框定)中,针对长文档中为分类任务提供必要上下文段落的挑战,提出一种基于实体导向的启发式上下文选择方法。该方法仅使用 XLM-RoBERTa 模型,在有限上下文窗口条件下,即可实现与采用更大生成式语言模型进行监督微调相当甚至更优的分类性能。结果表明,合理设计上下文选择策略能显著提升小模型在复杂文本分类任务中的表现。
原文摘要 · Abstract (English)
Our contribution to the SemEval 2025 shared task 10, subtask 1 on entity framing, tackles the challenge of providing the necessary segments from longer documents as context for classification with a masked language model. We show that a simple entity-oriented heuristics for context selection can enable text classification using models with limited context window. Our context selection approach and the XLM-RoBERTa language model is on par with, or outperforms, Supervised Fine-Tuning with larger generative language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。