为对话数据中的跨说话人依赖关系制定标注规范,支持协作构建等现象分析。
Coconstructions in spoken data: UD annotation guidelines and first results
- 提出基于说话人轮次和依赖关系的双轨标注框架。
- 区分重述与修正,并标记未完成短语中的关键成分。
- 适用于语音语料库中对话结构分析,适合语言学与NLP研究者。
本文为通用依存框架(Universal Dependencies)内的口语语料库,提出针对跨说话人轮次的句法依存标注指南,涵盖协作式共构、疑问句回答及回应性话语等现象。提出两种表示方式:一种是按说话人轮次分段的说话人导向表示,另一种是允许跨轮次依存的依赖关系表示。同时,提出了区分重述与修正的新标准,并建议突出未完成短语中的关键成分,以更好捕捉口语互动中的动态构建过程。
原文摘要 · Abstract (English)
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。