Avenzoar
The arcade

Casual on the surface.
Serious data collection underneath.

Every game is a data-collection task in disguise. Flip any card to see exactly what it collects and why a client wants it.

12 games and counting

Every game is
secretly collecting data.

These are formats we've already built. Drop your own data into one - or we design a brand-new game around your task. Hover any game to see what it collects.

Diacritization
Tashkeel
Tashkeel
Diacritization
flip
1
Diacritization

Tashkeel

Read the sentence, tap the right diacritic on each letter.

Written Arabic usually drops short vowels - so the same letters can be many different words. Restoring diacritics (tashkeel) is one of the hardest, highest-value problems in Arabic NLP.

text · MSA + dialects
OCR / handwriting validation
OCR Check
OCR Check
OCR / handwriting validation
flip
1
OCR / handwriting validation

OCR Check

See the original scan and the machine reading. Match or wrong?

Arabic OCR - especially handwriting and historical manuscripts - is notoriously unreliable. Human verification turns raw OCR into gold transcription data.

vision + text
Response quality rating
Answer Scale
Answer Scale
Response quality rating
flip
1
Response quality rating

Answer Scale

Weigh a question and its answer: excellent, partial, or wrong.

Rating answer quality is the backbone of RLHF and model evaluation. A balanced crowd of native speakers produces the preference data that aligns models.

preference · RLHF
Dialect identification
Dialect Compass
Dialect Compass
Dialect identification
flip
1
Dialect identification

Dialect Compass

Read the text, point the compass at its dialect.

Arabic is dozens of distinct spoken dialects, not one language. Labeled dialect data is scarce - and essential for models that understand how people actually talk.

Gulf · Egyptian · Levantine · Maghrebi
Content moderation
Red Card
Red Card
Content moderation
flip
1
Content moderation

Red Card

You're the referee. Clean, borderline, or a violation?

Safety classifiers need culturally-aware judgments of what counts as toxic or offensive - context that off-the-shelf models routinely get wrong.

safety · toxicity
Semantic similarity
Same Meaning?
Same Meaning?
Semantic similarity
flip
1
Semantic similarity

Same Meaning?

Two sentences appear. Identical, related, or unrelated?

Paraphrase and semantic-similarity labels train search, retrieval, and deduplication - and good similarity data is hard to come by.

STS · paraphrase
Named-entity recognition
Entity Radar
Entity Radar
Named-entity recognition
flip
1
Named-entity recognition

Entity Radar

Lock onto the highlighted term - person, place, org, or date?

Entity tagging powers everything from knowledge graphs to redaction. NER suffers from sparse, inconsistent labels - exactly what a big crowd fixes.

NER · 4 entity types
Fact verification
Trivia Runner
Trivia Runner
Fact verification
flip
1
Fact verification

Trivia Runner

Run the road, pick a lane: TRUE or FALSE.

True/false judgments at speed build fact-checking and claim-verification datasets - and keep players in flow while they label.

true / false
Cloze / text completion
Fill the Blank
Fill the Blank
Cloze / text completion
flip
1
Cloze / text completion

Fill the Blank

Spell the missing word on a scored keyboard.

Fill-in-the-blank responses generate completion and next-word data, and reveal how real speakers phrase things in context.

completion
Multiple-choice labeling
Swipe Quiz
Swipe Quiz
Multiple-choice labeling
flip
1
Multiple-choice labeling

Swipe Quiz

Flick toward your answer - four directions, four choices.

Fast multiple-choice rounds collect classification labels across any domain you can phrase as a question.

MCQ
Multiple-choice labeling
Bubble Quiz
Bubble Quiz
Multiple-choice labeling
flip
1
Multiple-choice labeling

Bubble Quiz

Pop the right bubble before the timer runs out.

A second MCQ format keeps the same labeling task fresh - variety is what keeps the crowd annotating for hours.

MCQ
Streak labeling
Stack-It
Stack-It
Streak labeling
flip
1
Streak labeling

Stack-It

Right answers stack the tower. Wrong ones topple it.

Endless streak play maximizes labels-per-session and naturally surfaces the questions players find hardest.

endless
New game formats ship regularly.
Any data a client needs - if we can turn it into a 30-second game, we can collect it.