corpus linguistics
Sign in to savebranch of linguistics that studies language through examples contained in real texts
Described at
Korpuslinguistik | Leuphana
leuphana.de →Link to a page describing this subject · 4,743 chars · not written by Vinony
Wikidata facts
- Image
- Lex Behatokia pantaila.png
Show 2 more facts
- Commons category
- Corpus linguistics
- described at URL
- www.leuphana.de/institute/ies/lina-lab/korpuslinguistik.html
via Wikidata · CC0
~12 min read
Article
Corpus linguistics is an empirical method for the study of language by text corpus (plural corpora). Corpora are balanced, often stratified collections of authentic, "real world", text of speech or writing that aim to represent a given linguistic variety. Today, corpora are generally machine-readable data collections.
Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. Large collections of text, though corpora may also be small in terms of running words, allow linguists to run quantitative analyses on linguistic concepts that may be difficult to test in a qualitative manner.