Three search targets
A token is one occurrence in a cited text. A normalized form is an editor’s standardized representation. A lemma is a conventional headword that groups forms of a lexeme. Searching the printed token, the normalized string and the lemma can produce different results. In Caesar’s line, the inflected divisa leads to the dictionary headword dīvidō. The dictionary entry even cites this Caesar passage, but its other senses and examples must not all be imported into this clause.
Coptic SCRIPTORIUM documents original spelling, normalized words and lemmas as distinct annotation layers. Its ANNIS tutorial shows how a lemma query can gather forms that an exact-spelling search would miss. That is useful for studying usage, but it depends on editorial segmentation and lemmatization. Record which layer you searched and which project’s decisions link it to the other layers.
A failed lookup
If a headword search returns nothing, inspect the original form and the project’s normalization policy before declaring it unknown. A damaged sign, variant spelling, abbreviation, homograph or mistaken word division can interrupt the link. A lemmatizer trained on one corpus may handle its conventions better than those of another. Search alternatives deliberately, retaining the unsuccessful query in your notes so that a later reader can reproduce the result.
When a form maps to two possible lemmas, do not choose by English gloss alone. Compare grammatical category, construction, period and cited examples. If both survive, keep both in the reading record.
Practice
In a documented corpus, search one occurrence by exact spelling and then by lemma. Write down the two query strings, filters and hit sets. Open two hits returned only by the lemma search and explain why their written forms differ.
Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.
