TrOCR Fine-tuned on Church Slavonic Handwriting (QuantiSlav)

This model is a fine-tuned version of kazars24/trocr-base-handwritten-ru for recognizing Church Slavonic handwritten text in Old Cyrillic script. It was trained by Achim Rabus (Slavic Department, University of Freiburg) on the dataset used for the Transkribus model generic-church-slavonic-handwriting-3.

Model Description

TrOCR (Transformer-based OCR) is a vision-to-text model using a ViT encoder and a causal language model decoder. This version is fine-tuned specifically on Church Slavonic handwriting in uncial and semi-uncial Old Cyrillic styles.

  • Base model: kazars24/trocr-base-handwritten-ru
  • Intended use: OCR/HTR of Church Slavonic historical manuscripts
  • Output: UTF-8 text including titlos, abbreviation marks, and diacritics

Training Data

Trained on Church Slavonic handwriting exported from Transkribus. The Transkribus model was trained by Elena Renje as part of the QuantiSlav project and curated by Achim Rabus (Slavic Department, University of Freiburg). The dataset covers Old Cyrillic script styles (uncial and semi-uncial), primarily East Slavic with South Slavic material included.

Main source manuscripts:

  • Codex Suprasliensis (10th–11th c., South Slavic recension)

  • Catecheses of Cyril of Jerusalem (transmitted version: 11th c., East Slavic recension)

  • Methodius of Olympus: Symposion (transmitted version: 17th c., East Slavic recension)

  • Velikie Minei Četʹi (16th c., East Slavic recension): large parts of the March and May volumes, Apostolos from the June volume

  • Size: 309,959 training lines (2,643 pages); 20,679 validation lines (205 pages)

  • Preprocessing: resize to 128 px height, aspect ratio preserved (LANCZOS); no background normalization

  • Validation: a subset of the validation lines was used for in-training evaluation

Performance

Metric Value
CER (validation) 4.99%

Note: A CNN + BiLSTM + CTC ("CRNN-CTC") model trained on the same QuantiSlav data reaches a lower CER on this validation set (2.89%, see achimrabus/crnn-ctc-church-slavonic). The two models are best seen as complementary — it is worth comparing both on your own material rather than relying on the validation CER alone.

Training Details

Parameter Value
Base model kazars24/trocr-base-handwritten-ru
Optimizer Adafactor
Learning rate 5e-5
Effective batch size 256 (per-device 64 × 4 GPUs, DDP)
Epochs ~7.4 (best checkpoint at step 9000; configured max 10)
FP16 Yes
Augmentation Rotation ±2°, brightness/contrast ±0.3
Generation max length 128
Framework HuggingFace Transformers, Seq2SeqTrainer

How to Use

from transformers import TrOCRProcessor, VisionEncoderDecoderModel
from PIL import Image

processor = TrOCRProcessor.from_pretrained("cyrillic-trocr/trocr-church-slavonic-handwritten")
model = VisionEncoderDecoderModel.from_pretrained("cyrillic-trocr/trocr-church-slavonic-handwritten")

# The model was trained on line images normalized to 128 px height with aspect
# ratio preserved. Reproduce that before the processor (which otherwise squashes
# directly to 384x384), to match the training distribution.
def resize_to_height(img, height=128):
    w, h = img.size
    return img.resize((max(1, round(w * height / h)), height), Image.LANCZOS)

image = resize_to_height(Image.open("line_image.png").convert("RGB"))
pixel_values = processor(images=image, return_tensors="pt").pixel_values

generated_ids = model.generate(pixel_values)
text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(text)

Note: Input should be a single text line image, not a full page.

Recommended: Polyscriptor

For real-world use on whole manuscript pages, this model is best run through Polyscriptor, a multi-engine HTR training and comparison tool. Polyscriptor handles automatic line segmentation, full-page batch processing, PAGE XML import/export, and lets you compare this TrOCR model directly against the CRNN-CTC model and other engines on your own material. It is available as a browser UI, a PyQt UI, and a command-line interface.

Intended Use

  • Transcription of Church Slavonic historical manuscripts
  • Research in Slavic medieval studies and digital humanities

Acknowledgements

Training data prepared within the QuantiSlav project (Elena Renje, Achim Rabus). Slavic Department, University of Freiburg.

Citation

TrOCR architecture:

@article{li2021trocr,
  title   = {TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models},
  author  = {Li, Minghao and Lv, Tengchao and Chen, Jingye and Cui, Lei and Lu, Yijuan and
             Florencio, Dinei and Zhang, Cha and Li, Zhoujun and Wei, Furu},
  journal = {arXiv preprint arXiv:2109.10282},
  year    = {2021}
}
Downloads last month
100
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cyrillic-trocr/trocr-church-slavonic-handwritten

Finetuned
(5)
this model

Paper for cyrillic-trocr/trocr-church-slavonic-handwritten