The research question
Begin with a question whose evidence can be observed: “How is this verb written in this group of letters?” or “Which forms of this lemma occur in this text?” Specify language, period, genre and the unit you will count. A corpus is a selected, edited body of material, not a complete record of a language. Its contents reflect survival, excavation, publication, digitization and editorial choices.
Write a one-line scope before searching: “Greek documentary papyri with transcriptions in the selected date range,” for example. A result outside that scope may be interesting, but it cannot quietly enter the same comparison.
Coverage and provenance
Read the project introduction and the description of each component. Papyri.info aggregates records from DDbDP, HGV, APIS and other partners: a catalogue record, transcription, translation and image may come from different contributors. Oracc is organized into separately managed projects, some with text corpora and some with portals only. A search across “Oracc” and a search inside one Oracc project therefore have different populations.
For each corpus, note which works or objects are included, which are excluded, the edition or transcription on which the searchable text rests, and whether annotations cover every item. Check when the corpus was last updated. If those facts are absent, treat the sample as exploratory.
A selection test
Compare two possible resources against the same question. The Digital Corpus of Sanskrit provides sandhi-split, morphologically and lexically analysed texts, making it useful for a lemma-and-form question. The Chinese Text Project offers searchable early Chinese texts and, for some passages, an electronic text aligned with a scanned source. Their search units and editorial layers differ; neither is a substitute for the other merely because both return text hits.
Choose the smallest collection that answers the question. Record the collection name, included texts, edition basis, annotation coverage and access date in a short source card.
Practice
Take one word or construction from a passage you are reading. Name two candidate corpora. For each, locate a statement of coverage and one sample record. State which searchable layer would answer your question, which layer is missing, and why you selected one corpus for the first search.
Modern-language examples on this page clarify a linguistic pattern. They do not establish unattested sounds or forms for an ancient language. Follow the cited sources below and keep editorial readings separate from the surviving trace.
