针对跨语言口语对话的语法解析难题,提出新基准与评估方法。
Lost in Speech: Benchmarking, Evaluation, and Parsing of Spoken Bilingual Conversational Language Beyond Standard UD Assumptions
- 构建英西双语口语对话的标注基准与现象分类体系
- 提出可区分严重错误与合理变异的评估指标Flex-UD
- 设计解耦式框架DECAP,提升复杂口语结构解析效果
口语双语对话因包含不连贯表达和受话语驱动的结构,给依赖句法解析带来挑战,尤其在标准通用依存(UD)假设下。为此,本文提出一个语言学基础的现象分类体系,并发布专家标注的英西双语口语对话基准SpokeBench。针对现有评估方法局限,提出灵活的模糊性感知评估指标Flex-UD,能区分灾难性结构失败与语言上可接受的变异。进一步提出DECAP解耦式智能解析框架,将口语现象处理与核心句法分析分离,实现无需重训练的鲁棒且可解释的依赖解析。在专有及开源大模型上的实验表明,DECAP在复杂口语现象上性能显著提升,相比基线在UPOS-F1得分上提高超60%,而传统基于连接的评估指标难以揭示此类增益。
原文摘要 · Abstract (English)
Spoken bilingual conversations pose substantial challenges for syntactic parsing because they often include disfluencies and discourse-driven structures that complicate dependency parsing under standard Universal Dependencies (UD) assumptions and evaluation practices. To systematically study these challenges, in this work, we first introduce a linguistically grounded taxonomy of conversational bilingual phenomena, together with SpokeBench, an expert-annotated English-Spanish benchmark for structurally complex speech. To address the limitations of existing evaluation practices, we propose Flex-UD, an ambiguity-aware evaluation metric that distinguishes catastrophic structural failures from linguistically acceptable variations. Finally, we introduce DECAP, a decoupled agentic parsing framework that separates spoken-phenomena handling from core syntactic analysis, enabling robust and interpretable dependency parsing without retraining. Experiments across both proprietary and open-weight LLMs show that DECAP substantially improves performance on complex conversational phenomena and achieves over 60% improvements in UPOS-F1 Score over baselines, while Flex-UD evaluations reveal gains that otherwise remain partially hidden under standard attachment-based metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。