通过词在上下文中的语义分析,提升对标注者分歧的建模能力。
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
- 融合多种相似性与距离度量,增强上下文嵌入表征。
- 使用适配器块提取任务特定表示,提升模型性能。
- 在分歧建模任务中表现良好,适合语义差异分析场景。
本文在CoMeDi共享任务中评估了词在上下文(WiC)任务中的标注者分歧,探究上下文语义与分歧之间的关系。以往研究多基于单句输入分析标注者属性,而本次任务引入WiC,旨在连接句子级语义表征与标注判断的变异性。我们提出了三种方法:一种特征增强方法,结合拼接、逐元素差值、乘积、余弦相似度及欧氏、曼哈顿距离扩展上下文嵌入;一种通过适配器块转换获得任务特定表示的方法;以及不同复杂度的分类器,包括集成模型。实验表明,包含丰富特征和任务特定表示的方法性能更优。尽管在子任务1(OGWiC)中未超越最优系统,但在子任务2(DisWiC)中表现接近官方评估结果,具备竞争力。
原文摘要 · Abstract (English)
In this work, we evaluate annotator disagreement in Word-in-Context (WiC) tasks exploring the relationship between contextual meaning and disagreement as part of the CoMeDi shared task competition. While prior studies have modeled disagreement by analyzing annotator attributes with single-sentence inputs, this shared task incorporates WiC to bridge the gap between sentence-level semantic representation and annotator judgment variability. We describe three different methods that we developed for the shared task, including a feature enrichment approach that combines concatenation, element-wise differences, products, and cosine similarity, Euclidean and Manhattan distances to extend contextual embedding representations, a transformation by Adapter blocks to obtain task-specific representations of contextual embeddings, and classifiers of varying complexities, including ensembles. The comparison of our methods demonstrates improved performance for methods that include enriched and task-specfic features. While the performance of our method falls short in comparison to the best system in subtask 1 (OGWiC), it is competitive to the official evaluation results in subtask 2 (DisWiC).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。