Language
Help

Corpora and search

What a corpus contains

Understand selection, editions and searchable layers.

Old Assyrian letter, The Met 66.245.1

The collection

A corpus is a collection assembled for a purpose. Its contents depend on what survived, what has been excavated, what editors published and what has been encoded. DCCLT focuses on lexical texts; its frequencies cannot be treated as a frequency profile of all Mesopotamian speech. Record the collection’s scope, version, periods and genres before drawing a general conclusion from it.

Search layers

A corpus may allow searches over diplomatic transcription, transliteration, normalized word forms, lemmas, translations or grammatical tags. These layers answer different questions. Searching a lemma can retrieve multiple spellings but may exclude uncertain tokens or include editorially restored ones. Keep the query and searched layer in your notes.

Damaged text

Oracc’s ATF notation distinguishes signs that are damaged, lost or read with uncertainty. Papyri.info likewise documents editorial signs. If a search interface strips these marks in snippets, open the full line and source image before counting a result as fully visible evidence.

Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.