corpus linguistics
Sign in to savebranch of linguistics that studies language through examples contained in real texts
In the Vinony graph
Within Vinony's link graph, corpus linguistics is referenced by 525 other articles, and connects out to linguistics, Sanskrit and dictionary.
Vinony files it under Applied linguistics, Corpus linguistics and Discourse analysis.
Its subject is documented across 43 Wikipedia language editions.
Described at
Korpuslinguistik | Leuphana
leuphana.de →Link to a page describing this subject · 4,743 chars · not written by Vinony
Wikidata facts
- Image
- Lex Behatokia pantaila.png
Show 2 more facts
- Commons category
- Corpus linguistics
- described at URL
- www.leuphana.de/institute/ies/lina-lab/korpuslinguistik.html
via Wikidata · CC0
~12 min read
Encyclopedic overview
Corpus linguistics is an empirical method for the study of language by text corpus (plural corpora). Corpora are balanced, often stratified collections of authentic, "real world", text of speech or writing that aim to represent a given linguistic variety. Today, corpora are generally machine-readable data collections.
Corpus linguistics proposes that a reliable analysis of a language is more feasible with corpora collected in the field—the natural context ("realia") of that language—with minimal experimental interference. Large collections of text, though corpora may also be small in terms of running words, allow linguists to run quantitative analyses on linguistic concepts that may be difficult to test in a qualitative manner.
Excerpted from Wikipedia’s “corpus linguistics” article, available under the CC BY-SA 4.0 licence.