Language
Help

Lexicography

Lemmatization and parsers

Use morphological tools as candidate generators and audit the decisions hidden in their output.

Field-dimension tablet from Tello, A31736

What a lemmatizer does

A lemmatizer maps tokens or word forms to dictionary entries. For an inflected language this usually requires a morphological analysis; for a script with ambiguous signs it may also require editorial readings and normalization. A parser can return part of speech, case, number, tense or other features, often several alternatives for one form. The output represents analyses permitted by the model or database. It does not itself establish which analysis the author intended in this sentence.

The ETCSL describes its own workflow in useful detail: it connects Sumerian forms to an ePSD citation form, part of speech, guide word and morphological analysis; some forms were not recognized automatically and required manual treatment. This is a concrete warning against treating a “no match” result as evidence that a word never existed. The Egyptian TLA connects text tokens with lemma lists, so a user can move between a corpus occurrence and the lexical record.

Ambiguity and errors

Homographs share a written form but may belong to different lexemes. Syncretic forms occupy several cells of one paradigm. A parser may therefore return multiple legitimate readings before context is consulted. It may also miss proper names, damaged passages, dialect spellings or rare constructions. Greek accents, Hebrew vowel signs, Coptic word segmentation and cuneiform transliteration conventions can change the search path. Record what string you entered, whether you stripped diacritics and which tool version produced the analysis.

Logeion explains that a fully written inflected form can suggest one or more lemmas, with hover information about parses. Its guidance also warns that transliterated Greek can be confused with literal Latin input if a suggested Greek form is not selected. Interface behavior is part of the evidence chain; a search that happened to land on a plausible entry has not validated the parse.

A manual check

Before accepting a parser result, identify the local syntactic role, agreement partners and plausible inflectional endings from a grammar. Then inspect the candidate lemma entry and cited uses. If two parses survive, retain both and state which evidence would distinguish them. A lemmatizer may be trained or hand-curated for one corpus, period or orthography; do not transfer its certainty to a different body of texts without testing.

Practice

Select one ambiguous form from a published text and run it through a documented word-study tool. List every returned lemma and parse. For each, write one reason the local grammar favors or rules it out. Preserve the original token and query. If the tool returns nothing, try a documented alternative spelling and report both searches.

One ETCSL line · form, lemma and gloss
Transliterated formETCSL lemma after reviewETCSL label
ugnim-eugnimtroops
igiigieye
im-ma-an-sig10sig10to place
ETCSL gives this line as an example of the lemmatization process. Its earlier automatic output grouped igi with a multiword lexeme; the reviewed analysis shown here treats the components separately. This difference is the lesson: a parser and a curated corpus can encode different analyses. University of Oxford, ETCSL: Lemmatisation ↗
Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.