Use cases

Your AI talks to thousands.
We read what it says.

Avenzoar puts a vetted crowd of human reviewers on the work machines can't grade on their own — auditing AI-agent conversations, evaluating model answers, checking Arabic transcripts and documents. Every judgment is cross-checked, disputes go to an expert, and you download findings your team can act on.

Jump to a use case
QA for AI agents & chatbots

Tens of thousands of conversations.
Nobody has read them.

Your agent answers customers around the clock. Your dashboards show how many conversations happened and how long they took — not whether the answers were right. Reading a few dozen by hand won't find the problems either: the failures that cost you hide in the long tail.

The problem

What hides in the long tail

Each of these is rare enough to slip past a spot check — and common enough to matter at your volume.

Wrong answers

Confident replies that are simply incorrect — the wrong price, the wrong procedure, the wrong eligibility rule.

Hallucinations

Policies, features and promises the agent invented because they sounded plausible.

Policy breaks

Replies that say, promise or disclose more than your agent is allowed to.

Tone that costs you

Curt, pushy or tone-deaf answers — often to customers who were already upset.

Missed escalations

Complaints, risks and requests a human should have taken over, left with the bot.

Dialect misunderstandings

A customer writes in Gulf, Levantine or Egyptian Arabic — and the agent answers a different question.

Full coverage or a smart sample

Review everything

Every conversation judged — for audits, launches and high-risk flows.

Review a designed sample

Sampled by topic, channel, dialect or date — or weighted toward conversations your own signals already flagged, like low ratings or repeat contacts.

How we do it

From raw logs
to clear verdicts.

01

Send the conversations

Export your logs as CSV or Excel. Share only the fields you're comfortable with, and keep your conversation and turn IDs — they come back with every label.

02

Agree the rubric

We turn your policies and quality bar into clear questions: Was it correct? Did it follow policy? Should it have escalated? Was the tone right?

03

Split into short review tasks

Each conversation — or each turn — becomes a quick judgment a trained reviewer makes on their phone, with the context they need right beside it.

04

Judge, check, settle

Several independent reviewers judge every item. Hidden attention checks and trust scores keep quality honest, and anything they disagree on goes to an expert for the final call.

05

Deliver the findings

Follow progress live in your client portal, then download the labeled data and a report your team can act on.

Your rubric, as short, clear questions

Every question is short, specific and tied to your policy — so answers stay consistent from one reviewer to the next.

  1. Did the agent understand what the customer asked?Yes / No
  2. Is the answer factually correct?Yes / No / Can't tell
  3. Did the reply stay within policy?Yes / No + which rule
  4. Should this have gone to a human?Yes / No
  5. How was the tone?1–5 rating
  6. What should the agent have said?Short written answer
Invented policy
Illustrative example
Customer
I was charged twice for the same order. Can I get the extra charge back?
Agent
Of course! Duplicate charges are refunded to your card instantly.
Reviewer verdictFail
Hallucination Missed escalation

Duplicate charges must be verified by the billing team before any refund. The agent promised an instant refund that policy doesn't allow.

Suggested fix

Never promise refund timing. For a duplicate charge, confirm the order number, then hand over to billing.

Dialect misread
Illustrative example
Customer
بدي ألغي الاشتراك، بس ما بدي يروح الرصيد اللي ضايل.
Translation: I want to cancel my subscription, but I don't want to lose my remaining balance.
Agent
لتحويل رصيدك إلى رقم آخر، اذهب إلى الإعدادات ثم «تحويل الرصيد».
Translation: To transfer your balance to another number, go to Settings, then "Transfer balance".
Reviewer verdictFail
Dialect misunderstanding Wrong answer

The customer asked how to cancel without losing their balance. The agent read it as a request to transfer balance.

Suggested fix

Add Levantine and Gulf phrasings of cancellation requests to the intent examples, and confirm intent when a message mentions both cancelling and balance.

iA made-up conversation to show the format — not client data.

What you get

Not just labels.
A list of what to fix.

A labeled dataset

Every conversation and turn with its verdicts, failure labels and level of agreement — in CSV or Excel, with your own IDs kept.

A failure taxonomy from your data

Which failures happen, how often and in which flows — counted on your conversations, not a generic benchmark.

The worst examples, first

The conversations most likely to hurt you rise to the top, so your team reads those before anything else.

Concrete fix suggestions

For each failure, reviewers write what the agent should have said. Grouped by pattern, those notes become specific prompt and guardrail changes.

A regression set

The failed conversations become a test set you replay after every fix — to confirm it held and nothing else broke.

Before / after comparison

Re-run the same review after your changes and see, category by category, what improved and what didn't.

Not a one-off audit

Find it. Fix it.
Prove it's fixed.

Your agent changes every time you touch a prompt, a model or a knowledge base. Run the review as a loop and every change ships with evidence.

1Review
People judge real conversations against your rubric.
2Diagnose
Failures are grouped, counted and ranked by impact.
3Fix
Apply prompt, knowledge-base and guardrail changes.
4Re-test
Replay the regression set and compare before / after.
Why Avenzoar

Human judgment,
without the headcount.

Fluent in Arabic dialects

Reviewers who read the Arabic your customers actually write — Gulf, Levantine, Egyptian and more — and notice when the agent didn't understand it.

Humans at scale, fast

Many reviewers work through your conversations in parallel, so tens of thousands of them don't wait on a small in-house team — and you never recruit, train or manage one.

Quality you can inspect

Consensus, attention checks, trust scores and expert adjudication — and the agreement level ships with every label.

Progress in plain sight

Follow the review in your client portal, explore items as they resolve, and download when you're ready.

Your IDs in, your IDs out

The columns you send come back in the export, so results join straight back to your logs, tickets and dashboards.

More use cases

Any human judgment,
at the scale you need.

The same crowd and the same quality checks, pointed at different problems. Each one runs on game formats we've already built.

LLM evaluation & preference ranking

Know which answer people actually prefer — and why.

The problem

Automated scores can't tell you which answer people actually prefer, or catch the subtle errors a native speaker spots at a glance. In Arabic, fluent-sounding and correct are often not the same thing.

How we do it

Reviewers compare two outputs side by side, rate answers against your rubric, or flag factual and safety errors. Several people judge each item and consensus settles the result.

What you get
  • Pairwise preferences ready for RLHF, DPO or reward models
  • Rubric scores per answer, with agreement
  • A clear winner per prompt when you compare models or prompt versions
Task formats A / B choice Rating True / false

Arabic speech & transcript QA

Transcripts checked by people who speak the dialect.

The problem

Speech models trained on formal Arabic stumble on dialects, code-switching and noisy calls — and a wrong transcript quietly corrupts everything built on top of it.

How we do it

Reviewers listen to each clip and approve or correct the machine transcript, or transcribe from scratch. When two versions compete, the crowd picks the right one.

What you get
  • Verified transcripts, checked word by word
  • Corrections matched to your clip IDs
  • A clean set to fine-tune or benchmark your speech model
Task formats Listen & correct Transcribe A / B choice

Document & OCR extraction checks

Every extracted field, confirmed by a human eye.

The problem

Handwriting, stamps, poor scans and connected Arabic script break extraction pipelines — silently, one field at a time.

How we do it

Reviewers see the image crop beside what your system extracted and confirm it or correct it — including handwritten text and equations.

What you get
  • Field-level corrections
  • Ground truth to measure and retrain your OCR
  • The error patterns your pipeline keeps repeating
Task formats Image & text check Correct the text Equation rebuild

Content moderation & safety labels

Your policy, applied by people who get the context.

The problem

Whether something is offensive depends on language, dialect and culture — and a policy written for English content rarely maps cleanly onto Arabic.

How we do it

Your policy becomes clear yes/no and category questions. Several reviewers judge each item, and borderline cases go to an expert.

What you get
  • Policy-aligned labels, with agreement
  • The borderline cases your policy doesn't cover yet
  • Training and evaluation data for safety classifiers
Task formats Clean or not Category True / false

Search & relevance judgments

Query by query: was the result actually useful?

The problem

You can't tune search, recommendations or RAG retrieval without knowing, query by query, whether the results were actually relevant.

How we do it

Reviewers grade query–result pairs, compare two rankings side by side, or mark whether two texts mean the same thing.

What you get
  • Graded relevance labels per query
  • Side-by-side verdicts between ranking versions
  • Evaluation sets for retrieval and RAG
Task formats Rating A / B choice Same or different

Arabic dialect & tashkeel data

Native-speaker data where it's scarcest.

The problem

Dialect-labeled and fully diacritized Arabic is scarce, and much of what exists was never checked by native speakers.

How we do it

Players identify the dialect of a sentence, add or verify diacritics, and tag people, places and organizations — every answer cross-checked by others.

What you get
  • Dialect labels per sentence
  • Verified diacritization (tashkeel)
  • Entity tags for Arabic NER
Task formats Dialect ID Tashkeel Entity tagging
How it works

From your data
to answers you trust.

The same four steps, whatever you bring us.

01

Tell us the job

Share a sample and what you want to learn. We design the rubric and the review task with you.

02

We turn it into short tasks

Your data becomes quick judgments vetted reviewers make on their phones, with the context right beside them.

03

The crowd judges, consensus decides

Several independent reviewers per item, hidden attention checks, trust scores, and an expert for anything disputed.

04

Follow along and download

Watch progress in your client portal, then download CSV or Excel with your own IDs, plus a PDF or PowerPoint report.

No sample metrics here
You won't find invented percentages on this page. Every number we give you is counted from your own data, with the agreement behind it.
Start with a sample

Tell us what you need checked.
We'll design the review.

Share a few example conversations or items and what you want to learn from them. We'll come back with a proposed rubric, how we'd run the review, timelines and pricing.