Language
Help

Writing systems

Unicode, fonts and transliteration

Preserve source signs while making editions searchable and displayable.

Palenque Palace and Temple of the Inscriptions

Character and glyph

Unicode assigns code points to abstract characters and defines how text can be exchanged. A font supplies visible glyphs; rendering software composes and positions them. A missing glyph is therefore not proof that a historical sign lacks a Unicode character, and a matching digital glyph is not proof that the ancient object had exactly that shape. Check the encoded character, script coverage and font separately. Keep the original image for all claims about particular strokes.

Combining marks can be separate characters. In Devanagari, the encoded sequence for कि follows logical order even though the dependent i sign appears visually before क. In Hebrew, points attach to consonant bases. Normalization may change equivalent code-point sequences; it does not determine the language or editorial reading. Search systems should document whether they ignore points, case or diacritics.

Transliteration and transcription

Transliteration maps written signs into another notation under a stated scheme. Transcription aims to represent a linguistic reading or pronunciation, and normalization may regularize spelling. Scholarly traditions use these labels in different ways, so the edition’s own key is decisive. In cuneiform, Oracc’s Latin letters, capitals, braces and damage symbols encode sign-by-sign editorial information. Stripping punctuation to make a smooth Latin word loses that information.

Store separate fields for object reference, diplomatic signs, transliteration, normalized form and translation. Use a language tag and explicit direction for mixed scripts. Search may index several fields, but the interface should say which one a hit came from. A transliteration hit is not the same as finding the exact historical sign form in a photograph.

Digital diagnosis

If कि displays with the i sign on the wrong side, first inspect the code-point order and the font’s shaping support. If Hebrew and Latin words jump around near a number, inspect direction markup and logical character order. If two visually identical accented strings do not match in a search, inspect normalization and the query rules. Each symptom points to a different technical layer; none licenses rewriting the source witness.

Evidence and comparison

Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.