TrOCR Fine-tuned on Church Slavonic Handwriting (QuantiSlav)
This model is a fine-tuned version of kazars24/trocr-base-handwritten-ru for recognizing Church Slavonic handwritten text in Old Cyrillic script. It was trained by Achim Rabus (Slavic Department, University of Freiburg) on the dataset used for the Transkribus model generic-church-slavonic-handwriting-3.
Model Description
TrOCR (Transformer-based OCR) is a vision-to-text model using a ViT encoder and a causal language model decoder. This version is fine-tuned specifically on Church Slavonic handwriting in uncial and semi-uncial Old Cyrillic styles.
- Base model: kazars24/trocr-base-handwritten-ru
- Intended use: OCR/HTR of Church Slavonic historical manuscripts
- Output: UTF-8 text including titlos, abbreviation marks, and diacritics
Training Data
Trained on Church Slavonic handwriting exported from Transkribus. The Transkribus model was trained by Elena Renje as part of the QuantiSlav project and curated by Achim Rabus (Slavic Department, University of Freiburg). The dataset covers Old Cyrillic script styles (uncial and semi-uncial), primarily East Slavic with South Slavic material included.
Main source manuscripts:
Codex Suprasliensis (10th–11th c., South Slavic recension)
Catecheses of Cyril of Jerusalem (transmitted version: 11th c., East Slavic recension)
Methodius of Olympus: Symposion (transmitted version: 17th c., East Slavic recension)
Velikie Minei Četʹi (16th c., East Slavic recension): large parts of the March and May volumes, Apostolos from the June volume
Size: 309,959 training lines (2,643 pages); 20,679 validation lines (205 pages)
Preprocessing: resize to 128 px height, aspect ratio preserved (LANCZOS); no background normalization
Validation: a subset of the validation lines was used for in-training evaluation
Performance
| Metric | Value |
|---|---|
| CER (validation) | 4.99% |
Note: A CNN + BiLSTM + CTC ("CRNN-CTC") model trained on the same QuantiSlav data reaches a lower CER on this validation set (2.89%, see
achimrabus/crnn-ctc-church-slavonic). The two models are best seen as complementary — it is worth comparing both on your own material rather than relying on the validation CER alone.
Training Details
| Parameter | Value |
|---|---|
| Base model | kazars24/trocr-base-handwritten-ru |
| Optimizer | Adafactor |
| Learning rate | 5e-5 |
| Effective batch size | 256 (per-device 64 × 4 GPUs, DDP) |
| Epochs | ~7.4 (best checkpoint at step 9000; configured max 10) |
| FP16 | Yes |
| Augmentation | Rotation ±2°, brightness/contrast ±0.3 |
| Generation max length | 128 |
| Framework | HuggingFace Transformers, Seq2SeqTrainer |
How to Use
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
from PIL import Image
processor = TrOCRProcessor.from_pretrained("cyrillic-trocr/trocr-church-slavonic-handwritten")
model = VisionEncoderDecoderModel.from_pretrained("cyrillic-trocr/trocr-church-slavonic-handwritten")
# The model was trained on line images normalized to 128 px height with aspect
# ratio preserved. Reproduce that before the processor (which otherwise squashes
# directly to 384x384), to match the training distribution.
def resize_to_height(img, height=128):
w, h = img.size
return img.resize((max(1, round(w * height / h)), height), Image.LANCZOS)
image = resize_to_height(Image.open("line_image.png").convert("RGB"))
pixel_values = processor(images=image, return_tensors="pt").pixel_values
generated_ids = model.generate(pixel_values)
text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(text)
Note: Input should be a single text line image, not a full page.
Recommended: Polyscriptor
For real-world use on whole manuscript pages, this model is best run through Polyscriptor, a multi-engine HTR training and comparison tool. Polyscriptor handles automatic line segmentation, full-page batch processing, PAGE XML import/export, and lets you compare this TrOCR model directly against the CRNN-CTC model and other engines on your own material. It is available as a browser UI, a PyQt UI, and a command-line interface.
Intended Use
- Transcription of Church Slavonic historical manuscripts
- Research in Slavic medieval studies and digital humanities
Acknowledgements
Training data prepared within the QuantiSlav project (Elena Renje, Achim Rabus). Slavic Department, University of Freiburg.
Citation
TrOCR architecture:
@article{li2021trocr,
title = {TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models},
author = {Li, Minghao and Lv, Tengchao and Chen, Jingye and Cui, Lei and Lu, Yijuan and
Florencio, Dinei and Zhang, Cha and Li, Zhoujun and Wei, Furu},
journal = {arXiv preprint arXiv:2109.10282},
year = {2021}
}
- Downloads last month
- 100
Model tree for cyrillic-trocr/trocr-church-slavonic-handwritten
Base model
microsoft/trocr-base-handwritten