Language
Help

Translation and tools

Digital tools and machine translation

Test tools against the edition, grammar and passage rather than their fluent output.

Field-dimension tablet from Tello, A31736

Match a tool to a question

A concordance locates occurrences; a lemmatizer proposes headwords; a morphological parser proposes features; an OCR or sign-recognition system proposes a transcription; machine translation proposes a target-language sentence. These are different operations and their errors compound. Coptic SCRIPTORIUM provides documented segmentation, normalization, tagging and lemmatization tools while also describing human correction of the resulting annotations. A tool output should be saved with the input string, model or project version, date and settings.

Ask what the tool was built to read. Script, dialect, period, genre, training corpus and editorial conventions all affect a result. Do not treat success on familiar classical prose as evidence of equal performance on an unpublished fragment. A confidence score, if offered, needs calibration against comparable held-out data before it can be interpreted as a probability that this passage is right.

Evaluation by error type

Machine-translation research distinguishes fluency from adequacy: English can read naturally while adding, omitting or mistranslating source content. WMT quality-estimation work explicitly examines such errors, including hallucinations. For study, make a small evaluation sheet: source span, tool output, edition-based analysis, omission or addition, lexical sense, grammatical role, uncertainty marker and severity. Check the source before comparing with a reference translation; two outputs that agree may share a borrowed interpretation.

Back-translation is a weak test because the second system can smooth the first system’s error. A more useful test asks the tool to expose the form, lemma, grammatical evidence and cited comparable usage, then independently checks each claim. If the system invents an unattested ancient form or a nonexistent citation, discard that part of the output.

Practice

Run a documented tool on the six-word Caesar line. Compare its English with your manual parse. Score whether it preserves the subject, the scope of omnis, the passive predicate and the number of parts. Record any addition, omission or unsupported certainty.

A source-based tool check
CheckQuestionEvidence to inspect
FormDid it read every source token?Edition and input string
GrammarDid it preserve subject, agreement and voice?Parser candidates and grammar
MeaningDid it add or omit a claim?Source clause and contextual lexicon
UncertaintyDid it erase a gap or restoration?Witness and edition symbols
The checklist is an original teaching adaptation of translation quality review, tailored to primary-source study. WMT 2023 Shared Task on Quality Estimation, Findings (ACL publication) ↗
Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.