SIMPLANG

Vocabulary Coverage Calculator

Every other tool here measures the effort — hours, streaks, words a day. This one measures what the effort buys. Enter the words you know and see how much of a text that actually covers, how many unknown words it leaves on a page, and what the next thousand is worth.

Word families, not individual forms — “run”, “runs” and “running” are one. If you only have a flashcard count, it is usually a little higher than your word-family count.
The baseline the coverage research is built on: general written text.

What that vocabulary is worth

You understand
87%
of the words in front of you
Unknown words per page (250 words)
33
about one in every 8
The next 1,000 words buy
4 points
of extra coverage
To 95% — adequate comprehension
5,000 words
3,000 to go
To 98% — unassisted reading
9,000 words
7,000 to go

Readable with a dictionary. You will follow the shape of a text but lose the detail, and guessing from context is still unreliable.

The next thousand words buy you 4 points of coverage. This is the steep part of the curve — nothing you learn later will pay like this.

Once you have a target worth aiming at, the vocabulary goal calculator turns it into a daily pace. English word-family figures from the Nation and Laufer coverage research; the curve transfers less cleanly to languages with heavier inflection or compounding.

Learning is linear. Understanding is not.

Pick a vocabulary target and most people reach for a round number. Ten thousand words sounds serious, so ten thousand it is. But a target only means something if you know what it buys, and the answer has been measured.

Word frequency is extremely uneven. The most common thousand word families cover about 80% of ordinary English text — one thousand words doing four fifths of the work. The next thousand adds around seven points. By the time you are at nine thousand, a further thousand words is worth about half a point. The effort per word never changes; the return collapses.

Two points on that curve have been tied to measured comprehension rather than intuition. At 95% coverage, Laufer found enough surrounding context that unknown words can be guessed — which is also the mechanism by which vocabulary grows from there. At 98%, Hu and Nation found readers could work unassisted. Between them lies most of the work of becoming a fluent reader, and it is worth knowing that before you start.

The most useful thing this calculator does is turn a percentage into something you can picture. Ninety-five percent sounds like nearly everything. Thirteen unknown words on every page does not.

Frequently Asked Questions

How many words do I need to understand a language?

It depends entirely on what you mean by understand, and the research gives two specific answers. At 95% coverage — about 5,000 word families in English — you have adequate comprehension: enough context around each unknown word to guess it, which is also how your vocabulary keeps growing. At 98%, around 8,000 to 9,000 families, you can read unassisted, without reaching for a dictionary. The gap between those two numbers is 4,000 words for three percentage points, which tells you something important about where the effort goes.

Why is 95% not nearly as good as it sounds?

Because 5% of a page is a lot of words. A typical page runs about 250 words, so at 95% coverage roughly 13 of them are unknown — one every two lines or so. That is enough to follow a text and still lose the argument, the joke, or the detail that mattered. It is why Hu and Nation's 1998 experiments found that readers needed 98% coverage before they could read comfortably without help: at that level the unknown words drop to about one in fifty, or five per page.

Why do the first thousand words matter so much more than the ninth?

Because word frequency is brutally uneven. A small set of words does most of the work in any language: the most frequent 1,000 word families cover roughly 80% of everything you read. The second thousand adds about 7 points, the third about 4, and by the ninth thousand a further thousand words is buying you half a percentage point. Learning stays linear — a thousand words is a thousand words — but understanding does not. That curve is the single most useful thing to know when deciding what to study next.

Should I stop learning words once the returns get small?

No, but you should change how. Once a further thousand words buys well under a point, general frequency lists are the wrong tool, because the words they give you next are ones you will rarely meet. What works from there is reading and listening to the things you actually care about and learning the vocabulary you meet in them — that vocabulary is low-frequency in general, but high-frequency in your life. The low-frequency tail is also where the specialist and technical words live, and those you learn from the field, not from a list.

What is a word family, and why not just count words?

A word family groups a base word with its inflections and regular derivations — run, runs, ran, running are one family, not four. The coverage research counts families because once you know the base and the pattern, the rest come almost free. This matters when you compare your own count to these figures: flashcard decks usually count individual forms, so your deck number tends to be higher than your word-family number. If your app reports word forms, expect your true family count to be somewhat lower.

Do these figures work for any language?

The shape does; the exact numbers do not. Frequency distributions are steep in every language, so the first thousand words always do the heavy lifting and the returns always diminish. But the figures here come from English word-family research using the British National Corpus, and languages with heavier inflection or productive compounding behave differently — a highly inflected language spreads the same meaning across more forms, and a compounding language lets you decode words you have never seen. Treat the percentages as the right shape and the right order of magnitude rather than a measurement of your language.