Language
Help

Lexicography

Tokens, forms and lexemes

Learn what a dictionary groups together and why the word on the object is not necessarily its headword.

Field-dimension tablet from Tello, A31736

The three units

A token is one occurrence in a particular text. A word form is the spelling or grammatical shape represented by that token. A lexeme is the vocabulary item under which related forms are grouped. A dictionary normally gives a lemma, or citation form, to represent the lexeme. The lemma is an access point created by lexicographical convention; it need not have appeared on the object you are reading. Ten occurrences of one form are ten tokens, not ten separate dictionary words.

This distinction immediately changes how you use a digital search. A token search asks where this written sequence occurs. A lemma search groups forms that editors or an algorithm have assigned to one lexeme. A full-text dictionary search may instead find a string anywhere in definitions and quotations. Ask which of these the search box actually searches before interpreting the result.

Finding a citation form

Perseus notes that an inflected Greek or Latin word in a text may not resemble the form under which a dictionary files it. Its Word Study Tool proposes analyses and dictionary headwords. That proposal is a bridge from token to lexeme; grammar and context must still decide among alternatives. The Oxford ETCSL similarly links Sumerian forms to an ePSD citation form, gloss and morphological analysis. Its lemmatisation account explicitly includes some multiword units. There is no universal rule that a lemma must be one graphic word.

A citation form may be a nominative singular noun, a first-person verb, an infinitive, a root, or another conventional representative according to the language and lexicon. Learn the dictionary preface and alphabetization rules. Do not silently replace an edition’s spelling with the headword in a quotation: preserve both.

Normalization

A normalized form standardizes some features of an attested spelling. It may expand an abbreviation, regularize a historical orthography or choose one transliteration of a polyvalent sign. Normalization can make searching possible, but it adds an editorial layer. The searchable token, normalized form and lemma should remain separable so another reader can reconstruct how an entry was reached. An uncertain sign cannot be made certain by the existence of a convenient dictionary headword.

Practice

Take three tokens from one edition. For each, record the edition and location, the exact printed token, any diplomatic reading, the proposed lemma and the rule or tool that connects them. Search the token and lemma separately. If they produce different sets of hits, explain the difference before choosing a translation.

Three layers of a lookup
LayerQuestionKeep in a reading note
TokenWhat appears at this location?Witness, line and exact transcription
Word formWhat grammatical or orthographic shape is this?Inflection, spelling and uncertainty
Lexeme / lemmaWhich vocabulary item could it represent?Dictionary, edition and candidate headwords
The distinctions are explanatory; no ancient form has been supplied or reconstructed in this table. Text Encoding Initiative, P5 Guidelines §10: Dictionaries ↗
Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.