What is Corpus Linguistics?
Corpus Linguistics is the analysis of collections of language, like that found in books, newspapers, or other 'real world' contexts. It also includes spoken language which has been meticulously recorded and transcribed by a number of people.
There are a number of corpora (the plural of corpus) available to use by linguists and other researchers.
British National Corpus (BNC):
This contains an astonishing 87,284,364 written words and 10,341,729 spoken words, spanning categories such as "informative writing" e.g. world affairs, arts, belief and thought, social/natural sciences etc., "imaginative writing" (fiction), "spoken context governed" (speech recorded at public talks), and "spoken demographic" (informal spoken speech).
These may be collections of English and Spanish words, or American English and Indian English. I imagine these may possibly be good for metaphor analysis. I'll have to find out more though.
These are corpora where exactly the same texts have been translated, for example, the CRATER corpus. I think these must be for translations and language-learning, but I've asked if there are additional usages of this type of corpus in the course forum.
This is language use created by people learning a particular language. An example of these is the International Corpus of Learner English. I guess this could be interesting if you want to find patterns in the way people learn a language.
Historical or Diachronic Corpus
This type is probably the most interesting to me. These corpora contain texts from a particular time period. The Helsinki Corpus, for example, contains 1.5 million words of texts from 700AD to 1700AD.
This is a very interesting type of corpus that is updated on a very regular basis. By doing so, we can see how language changes on a daily basis and can find the origins of new words in a language. An example of this is the Bank of English which is maintained by Birmingham University. Text is added to this as regularly as on a daily basis.