Category Archives: Access

Correct citation of COW corpora

If you use the COW corpora in publications (including talks and presentations) you must cite a paper by the corpus creators. Which paper you have to cite must be determined at the time of the submission of your publication by visiting the COW citation link.

Read this: You can only query and download sentence shuffles!

As explained in Section 4.1 of Roland Schäfer (2015) Processing and querying large web corpora with the COW14 architecture, we have to take certain measures in order to stay within the bounds of German copyright laws. This means that we only release sentence shuffles, i.e., corpora which are just bags of sentences. In other words, there are no documents in released versions of COW corpora, just single sentences without contexts. The original URL plus some other meta data are recorded for each sentence, however.

Downloads on webcorpora.org

The unified access at webcorpora.org will in the future offer downloads of sentence shuffles of COW corpora. It supersedes the legacy CODS (COW download system) hosted by the German Grammar Group at Freie Universität Berlin.