Mission
Why this exists
A pecha (པེ་ཆ་) is a traditional Tibetan book — long, loose leaves of scripture, copied and recopied by hand for a thousand years. Pecha Lab is an independent, non-commercial initiative that carries that work into open data: complete, validated, machine-readable editions of the Tibetan Buddhist canon, freely usable by anyone.
The texts belong to all Tibetan Buddhist traditions equally. We serve two audiences at once: communities and scholars who need trustworthy digital access to the canon, and researchers building natural-language-processing tools for Tibetan — a language of millions of speakers that remains severely under-resourced in modern computing. Careful open data advances both preservation and research.
The Dataset
The complete Derge canon, in Unicode
Our first release is a validated Unicode corpus of the complete Derge (sde dge) edition — both the Kangyur (the Buddha-word) and the Tengyur (the treatises) — built from the Buddhist Digital Resource Center’s public catalogue and etext archive, segmented into sentence-like units ready for research use.
Kangyur: 1,114 of 1,114 works — including 100 works recovered from a clean parallel digitization of the same edition after our validation gate flagged legacy-font corruption. Tengyur: 3,370 works. Every published text passes strict Unicode validation.
Method
How it is built
- A reproducible pipeline over BDRC’s public graph. The corpus is assembled by programmatically walking the Buddhist Digital Resource Center’s open RDF catalogue from each canonical root down to every electronic text — every step scripted, auditable, and re-runnable end to end.
- A corruption-detection validation gate. Legacy-font conversions can silently mis-decode Tibetan into other Indic scripts. Our gate measures the fraction of non-Tibetan Indic codepoints in each text, catching this corruption reliably while correctly retaining mantra-dense passages that simpler checks reject.
- Cross-edition recovery for a complete Kangyur. The 100 corrupted Kangyur works were recovered from a clean parallel digitization of the same Derge edition, matched text-by-text through BDRC’s abstract-work records — so the published Kangyur is complete, not merely filtered.
Roadmap
What comes next
- OCR of image-only witnesses — extending coverage to canonical texts that today exist only as page scans.
- Tibetan–English alignment — pairing canonical Tibetan with published English translations to support translation research.
- Multilingual title linking — connecting each work’s Tibetan title to its Sanskrit, English, and Chinese counterparts across the major catalogues.
Code & Data
Where to find it
The pipeline code and the dataset are being prepared for public release under the Pecha Lab organization:
Contact
Get in touch
Questions, corrections, or collaboration ideas: hello@pechalab.com