Avenzoar

BEFORE / AFTER

Same model.Different data.

Arabic breaks models trained on generic data — handwriting, dialects, diacritics. Flip the switch on each demo and watch what training on Avenzoar's human-verified data changes.

Interactive simulation for illustration — sample images are real items from our annotation games.

OCR · HANDWRITING

Reading real handwriting

A real item from our OCR game, exactly as it resolved: the machine read it without the final letter — every single player caught it. Same base model, different training data.

Base model: Gemini 3 Flash
Arabic signage sample

MODEL OUTPUT

بالكاميرا
38%character error rate (simulated)

HANDWRITTEN MATH

Equations, symbol by symbol

Exponents, variables and operators from a worksheet — rebuilt correctly only when the model has seen thousands of human-verified examples.

Base model: GPT-5
handwritten equation

MODEL OUTPUT

8xY+7x'l=O
52%expression error rate (simulated)

AUDIO · DIALECTS

Hearing the street, not the textbook

A real 10-second clip from our lesson-audio game — a Jordanian physics teacher. Off: the classic mistakes MSA-trained ASR makes. On: the transcript our players verified, word by word.

Base model: Whisper large-v3
Dialect: unknown

MODEL TRANSCRIPT

طيبالقوةالمركبةكيفبيديأحسبها؟أحيانابنعرفحسبقانوننيوتن

44%word error rate (simulated)

Bring your model. We'll bring the data.

Tell us where your model breaks — we'll build the game that fixes it.

Talk to us