整理80所高校的生成式AI使用指南,构建学术领域通用数据集。
AGGA: A Dataset of Academic Guidelines for Generative AI and Large Language Models
- 从全球6大洲高校官网收集80份AI使用规范文本
- 数据总量18.8万词,涵盖人文、科技等多学科场景
- 可作标注基准,支持歧义检测与需求分类等任务
本研究提出AGGA,一个包含80份学术机构生成式AI(GAIs)和大型语言模型(LLMs)使用指南的数据集,内容源自各大学官方网页。数据集总字数达188,674词,可用于自然语言处理中的模型合成、抽象识别与文档结构评估等需求工程任务。通过严谨方法选取覆盖六大洲的顶尖高校,涵盖人文、科技等领域,兼顾公立与私立机构,呈现多元视角。该数据集可进一步标注,用作歧义检测、需求分类及等价需求识别等任务的基准。
原文摘要 · Abstract (English)
This study introduces AGGA, a dataset comprising 80 academic guidelines for the use of Generative AIs (GAIs) and Large Language Models (LLMs) in academic settings, meticulously collected from official university websites. The dataset contains 188,674 words and serves as a valuable resource for natural language processing tasks commonly applied in requirements engineering, such as model synthesis, abstraction identification, and document structure assessment. Additionally, AGGA can be further annotated to function as a benchmark for various tasks, including ambiguity detection, requirements categorization, and the identification of equivalent requirements. Our methodologically rigorous approach ensured a thorough examination, with a selection of universities that represent a diverse range of global institutions, including top-ranked universities across six continents. The dataset captures perspectives from a variety of academic fields, including humanities, technology, and both public and private institutions, offering a broad spectrum of insights into the integration of GAIs and LLMs in academia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。