* * * * * * * * * * * * * * * * * * * * * * * *
LINGUIST List logo Eastern Michigan University Wayne State University *
* People & Organizations * Jobs * Calls & Conferences * Publications * Language Resources * Text & Computer Tools * Teaching & Learning * Mailing Lists * Search *
* *
LINGUIST List 20.284

Thu Jan 29 2009

Qs: Statistics of English Vocabulary

Editor for this issue: Dan Parker <danlinguistlist.org>

We'd like to remind readers that the responses to queries are usually best posted to the individual asking the question. That individual is then strongly encouraged to post a summary to the list. This policy was instituted to help control the huge volume of mail on LINGUIST; so we would appreciate your cooperating with it whenever it seems appropriate.

In addition to posting a summary, we'd like to remind people that it is usually a good idea to personally thank those individuals who have taken the trouble to respond to the query.

To post to LINGUIST, use our convenient web form at http://linguistlist.org/LL/posttolinguist.html.
        1.    Richard Hudson, Statistics of English Vocabulary

Message 1: Statistics of English Vocabulary
Date: 28-Jan-2009
From: Richard Hudson <dickling.ucl.ac.uk>
Subject: Statistics of English Vocabulary
E-mail this message to a friend

Dear All,

I wonder if someone could help me with two statistical question about the
vocabulary of English (as found in corpus work - at this point I'm not
asking for figures for individual speakers, though they would be really
fascinating to know if anyone has them).

Q1. How many morphemes are there? (I'm sure I've seen a figure somewhere,
the point being, of course, that it's much smaller than the number of
lexemes (lemmas, lexical items).

Q2. What percentage of the total vocabulary belongs to the various major
word classes? Better still, how does this percentage vary with frequency?
(I assume for example that rare words tend to be nouns.)

If there's enough response I'll summarise back to the list.

Best wishes, Dick Hudson

Linguistic Field(s): Text/Corpus Linguistics

Read more issues|LINGUIST home page|Top of issue

Please report any bad links or misclassified data

LINGUIST Homepage | Read LINGUIST | Contact us

NSF Logo

While the LINGUIST List makes every effort to ensure the linguistic relevance of sites listed
on its pages, it cannot vouch for their contents.