构建国会听证问答数据集,揭示两党提问策略差异。
C-QUERI: Congressional Questions, Exchanges, and Responses in Institutions Dataset
- 从听证会文本提取问答对,构建跨届次国会数据集
- 仅凭问题内容即可准确预测提问者党派归属
- 适用于政治传播、话语分析与对话研究者
政治访谈与听证会中的提问不仅用于获取信息,更具有推动党派叙事和塑造公众认知的战略目的。然而由于缺乏大规模语料库,此类话语的策略性特征长期未被充分研究。国会听证会因其制度化流程、强制回应机制及两党代表均有提问机会,成为研究政治提问的理想场景。本文开发了一套从非结构化听证记录中提取问答对的流水线,构建了第108至117届国会委员会听证会的新型数据集。分析显示,不同政党提问者存在系统性策略差异,仅凭问题内容即可准确预测提问者党派。该数据集与方法不仅推动了国会政治研究,也为各类问答式对话分析提供了通用框架。
原文摘要 · Abstract (English)
Questions in political interviews and hearings serve strategic purposes beyond information gathering including advancing partisan narratives and shaping public perceptions. However, these strategic aspects remain understudied due to the lack of large-scale datasets for studying such discourse. Congressional hearings provide an especially rich and tractable site for studying political questioning: Interactions are structured by formal rules, witnesses are obliged to respond, and members with different political affiliations are guaranteed opportunities to ask questions, enabling comparisons of behaviors across the political spectrum. We develop a pipeline to extract question-answer pairs from unstructured hearing transcripts and construct a novel dataset of committee hearings from the 108th--117th Congress. Our analysis reveals systematic differences in questioning strategies across parties, by showing the party affiliation of questioners can be predicted from their questions alone. Our dataset and methods not only advance the study of congressional politics, but also provide a general framework for analyzing question-answering across interview-like settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。