Building a DDC-annotated Corpus from OAI Metadata

Authors

Mathias Lösch Bielefeld University Library
Ulli Waltinger Bielefeld University
Wolfram Horstmann Bielefeld University Library
Alexander Mehler Frankfurt University

Abstract

Document servers complying to the standards of the Open Archives Initiative (OAI) are rich, yet seldom exploited source of textual primary data for research fields in text mining, natural language processing or computational linguistics. We present a bilingual (English and German) text corpus consisting of bibliographic OAI records and the associated full texts. A particular added value is that we annotated each record with at least one Dewey Decimal Classification (DDC) number, inducing a subject-based categorization of the corpus. By this means, it can be used as training data for machine learning-based text categorization tasks in digital libraries, but also as primary data source for linguistic research on academic language use related to specific disciplines. We describe the construction of the corpus using data from the Bielefeld Academic Search Engine (BASE), as well as its characteristics.

Downloads

Published

2011-04-29

Issue

Vol. 12 No. 2 (2011): Open Repositories 2010

Section

Articles