Language
Help

Corpora and search

Hits and frequencies

Turn a result list into a qualified observation with a meaningful denominator.

Old Assyrian letter, The Met 66.245.1

Count the right unit

A search result can count matching tokens, lines, documents or records. Repeated lines in copies of the same composition may produce many hits without many independent contexts. A composite edition may add another representation of the same passage. Open enough results to learn what one row means and decide whether the question is about occurrences, distinct texts or independent witnesses.

Write “12 matching records” when that is what the interface counted. Do not silently convert records to tokens. A page of results is also not a random sample of the corpus unless the sorting and pagination support that inference.

Denominators

A raw count of 20 versus 10 says little if one group contains ten times as much searchable text. Divide by the number of eligible tokens when tokenization is consistent, or by eligible documents for a document-level question. For example, 20 hits in 40,000 eligible tokens equals 5 per 10,000; 10 hits in 5,000 equals 20 per 10,000. This arithmetic teaches the comparison, but a rate is only as sound as its numerator, denominator and corpus selection.

Perseus's Vocabulary Tool describes frequencies for a selected set of works or sections. DCS detail pages report distributions for lexical units. Read which texts and forms their totals cover before carrying the numbers into your own table.

Bias and uncertainty

Publication and digitization favor some places, periods, genres and material types. Damage can remove searchable words. Normalization can make variants visible while obscuring their original spelling. Manual and automated tagging have error, and editorial revisions can change totals. A zero hit means the chosen query matched nothing in the indexed scope; it cannot establish that ancient speakers never used a form.

For a small study, report a simple sensitivity check: count exact written forms alone, then count the broader annotated or normalized set, and say why the results differ. Avoid percentages with false precision when many hits are ambiguous.

Practice

Make a two-row comparison between two defined subsets. Write the unit of the numerator, the denominator and the inclusion rule above the numbers. Inspect five hits from each subset and classify genuine, ambiguous and false hits. Recalculate a range if ambiguous examples materially change the comparison.

Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.