Avenzoar puts a vetted crowd of human reviewers on the work machines can't grade on their own — auditing AI-agent conversations, evaluating model answers, checking Arabic transcripts and documents. Every judgment is cross-checked, disputes go to an expert, and you download findings your team can act on.
Agent conversation reviewIllustrative
Reviewed
ConversationIn review Fail
I was charged twice for the same order. Can I get the extra charge back?
Of course! Duplicate charges are refunded to your card instantly.
Hallucination Missed escalation
Perfect, thank you!
Rubric
Understood the askCorrect & on policyRight toneEscalated when needed
Reviewers
Consensus
ConversationIn review Fail
بدي ألغي الاشتراك، بس ما بدي يروح الرصيد اللي ضايل.
I want to cancel, but I don't want to lose my remaining balance.
لتحويل رصيدك إلى رقم آخر، اذهب إلى الإعدادات ثم «تحويل الرصيد».
To transfer your balance to another number, go to Settings.
Dialect misread Wrong answer
لا، أنا بدي ألغي!
No, I want to cancel!
Rubric
Understood the askCorrect & on policyRight toneEscalated when needed
Reviewers
Consensus
ConversationIn review Fail
I've asked about my order three times now. Nobody is helping me.
As I already said, please check the tracking page.
Tone Missed escalation
Can I talk to a person, please?
Rubric
Understood the askCorrect & on policyRight toneEscalated when needed
Reviewers
Consensus
Regression set
The same conversations, re-tested
I was charged twice for the same order. Can I get the extra charge back?HallucinationMissed escalation
Illustrative — your report counts these on your own conversations.
Prompt fixSystem prompt
−Reassure customers that refunds happen right away.
+Never promise refund timing. Hand duplicate charges to billing.
+If a message mixes cancelling and balance, confirm the intent first.
+If a customer repeats a complaint, apologise and offer a person.
Jump to a use case
QA for AI agents & chatbots
Tens of thousands of conversations. Nobody has read them.
Your agent answers customers around the clock. Your dashboards show how many conversations happened and how long they took — not whether the answers were right. Reading a few dozen by hand won't find the problems either: the failures that cost you hide in the long tail.
The problem
What hides in the long tail
Each of these is rare enough to slip past a spot check — and common enough to matter at your volume.
Wrong answers
Confident replies that are simply incorrect — the wrong price, the wrong procedure, the wrong eligibility rule.
Hallucinations
Policies, features and promises the agent invented because they sounded plausible.
Policy breaks
Replies that say, promise or disclose more than your agent is allowed to.
Tone that costs you
Curt, pushy or tone-deaf answers — often to customers who were already upset.
Missed escalations
Complaints, risks and requests a human should have taken over, left with the bot.
Dialect misunderstandings
A customer writes in Gulf, Levantine or Egyptian Arabic — and the agent answers a different question.
Full coverage or a smart sample
Review everything
Every conversation judged — for audits, launches and high-risk flows.
Review a designed sample
Sampled by topic, channel, dialect or date — or weighted toward conversations your own signals already flagged, like low ratings or repeat contacts.
How we do it
From raw logs to clear verdicts.
01
Send the conversations
Export your logs as CSV or Excel. Share only the fields you're comfortable with, and keep your conversation and turn IDs — they come back with every label.
02
Agree the rubric
We turn your policies and quality bar into clear questions: Was it correct? Did it follow policy? Should it have escalated? Was the tone right?
03
Split into short review tasks
Each conversation — or each turn — becomes a quick judgment a trained reviewer makes on their phone, with the context they need right beside it.
04
Judge, check, settle
Several independent reviewers judge every item. Hidden attention checks and trust scores keep quality honest, and anything they disagree on goes to an expert for the final call.
05
Deliver the findings
Follow progress live in your client portal, then download the labeled data and a report your team can act on.
Your rubric, as short, clear questions
Every question is short, specific and tied to your policy — so answers stay consistent from one reviewer to the next.
Did the agent understand what the customer asked?Yes / No
Is the answer factually correct?Yes / No / Can't tell
Did the reply stay within policy?Yes / No + which rule
Should this have gone to a human?Yes / No
How was the tone?1–5 rating
What should the agent have said?Short written answer
Invented policy
Illustrative example
Customer
I was charged twice for the same order. Can I get the extra charge back?
Agent
Of course! Duplicate charges are refunded to your card instantly.
Reviewer verdictFail
Hallucination Missed escalation
Duplicate charges must be verified by the billing team before any refund. The agent promised an instant refund that policy doesn't allow.
Suggested fix
Never promise refund timing. For a duplicate charge, confirm the order number, then hand over to billing.
Dialect misread
Illustrative example
Customer
بدي ألغي الاشتراك، بس ما بدي يروح الرصيد اللي ضايل.
Translation: I want to cancel my subscription, but I don't want to lose my remaining balance.
Agent
لتحويل رصيدك إلى رقم آخر، اذهب إلى الإعدادات ثم «تحويل الرصيد».
Translation: To transfer your balance to another number, go to Settings, then "Transfer balance".
Reviewer verdictFail
Dialect misunderstanding Wrong answer
The customer asked how to cancel without losing their balance. The agent read it as a request to transfer balance.
Suggested fix
Add Levantine and Gulf phrasings of cancellation requests to the intent examples, and confirm intent when a message mentions both cancelling and balance.
iA made-up conversation to show the format — not client data.
What you get
Not just labels. A list of what to fix.
A labeled dataset
Every conversation and turn with its verdicts, failure labels and level of agreement — in CSV or Excel, with your own IDs kept.
A failure taxonomy from your data
Which failures happen, how often and in which flows — counted on your conversations, not a generic benchmark.
The worst examples, first
The conversations most likely to hurt you rise to the top, so your team reads those before anything else.
Concrete fix suggestions
For each failure, reviewers write what the agent should have said. Grouped by pattern, those notes become specific prompt and guardrail changes.
A regression set
The failed conversations become a test set you replay after every fix — to confirm it held and nothing else broke.
Before / after comparison
Re-run the same review after your changes and see, category by category, what improved and what didn't.
Not a one-off audit
Find it. Fix it. Prove it's fixed.
Your agent changes every time you touch a prompt, a model or a knowledge base. Run the review as a loop and every change ships with evidence.
1Review
People judge real conversations against your rubric.
2Diagnose
Failures are grouped, counted and ranked by impact.
3Fix
Apply prompt, knowledge-base and guardrail changes.
4Re-test
Replay the regression set and compare before / after.
Why Avenzoar
Human judgment, without the headcount.
Fluent in Arabic dialects
Reviewers who read the Arabic your customers actually write — Gulf, Levantine, Egyptian and more — and notice when the agent didn't understand it.
Humans at scale, fast
Many reviewers work through your conversations in parallel, so tens of thousands of them don't wait on a small in-house team — and you never recruit, train or manage one.
Quality you can inspect
Consensus, attention checks, trust scores and expert adjudication — and the agreement level ships with every label.
Progress in plain sight
Follow the review in your client portal, explore items as they resolve, and download when you're ready.
Your IDs in, your IDs out
The columns you send come back in the export, so results join straight back to your logs, tickets and dashboards.
More use cases
Any human judgment, at the scale you need.
The same crowd and the same quality checks, pointed at different problems. Each one runs on game formats we've already built.
LLM evaluation & preference ranking
Know which answer people actually prefer — and why.
The problem
Automated scores can't tell you which answer people actually prefer, or catch the subtle errors a native speaker spots at a glance. In Arabic, fluent-sounding and correct are often not the same thing.
How we do it
Reviewers compare two outputs side by side, rate answers against your rubric, or flag factual and safety errors. Several people judge each item and consensus settles the result.
What you get
Pairwise preferences ready for RLHF, DPO or reward models
Rubric scores per answer, with agreement
A clear winner per prompt when you compare models or prompt versions
Task formats A / B choice Rating True / false
Arabic speech & transcript QA
Transcripts checked by people who speak the dialect.
The problem
Speech models trained on formal Arabic stumble on dialects, code-switching and noisy calls — and a wrong transcript quietly corrupts everything built on top of it.
How we do it
Reviewers listen to each clip and approve or correct the machine transcript, or transcribe from scratch. When two versions compete, the crowd picks the right one.
What you get
Verified transcripts, checked word by word
Corrections matched to your clip IDs
A clean set to fine-tune or benchmark your speech model
Task formats Listen & correct Transcribe A / B choice
Document & OCR extraction checks
Every extracted field, confirmed by a human eye.
The problem
Handwriting, stamps, poor scans and connected Arabic script break extraction pipelines — silently, one field at a time.
How we do it
Reviewers see the image crop beside what your system extracted and confirm it or correct it — including handwritten text and equations.
What you get
Field-level corrections
Ground truth to measure and retrain your OCR
The error patterns your pipeline keeps repeating
Task formats Image & text check Correct the text Equation rebuild
Content moderation & safety labels
Your policy, applied by people who get the context.
The problem
Whether something is offensive depends on language, dialect and culture — and a policy written for English content rarely maps cleanly onto Arabic.
How we do it
Your policy becomes clear yes/no and category questions. Several reviewers judge each item, and borderline cases go to an expert.
What you get
Policy-aligned labels, with agreement
The borderline cases your policy doesn't cover yet
Training and evaluation data for safety classifiers
Task formats Clean or not Category True / false
Search & relevance judgments
Query by query: was the result actually useful?
The problem
You can't tune search, recommendations or RAG retrieval without knowing, query by query, whether the results were actually relevant.
How we do it
Reviewers grade query–result pairs, compare two rankings side by side, or mark whether two texts mean the same thing.
What you get
Graded relevance labels per query
Side-by-side verdicts between ranking versions
Evaluation sets for retrieval and RAG
Task formats Rating A / B choice Same or different
Arabic dialect & tashkeel data
Native-speaker data where it's scarcest.
The problem
Dialect-labeled and fully diacritized Arabic is scarce, and much of what exists was never checked by native speakers.
How we do it
Players identify the dialect of a sentence, add or verify diacritics, and tag people, places and organizations — every answer cross-checked by others.
What you get
Dialect labels per sentence
Verified diacritization (tashkeel)
Entity tags for Arabic NER
Task formats Dialect ID Tashkeel Entity tagging
How it works
From your data to answers you trust.
The same four steps, whatever you bring us.
01
Tell us the job
Share a sample and what you want to learn. We design the rubric and the review task with you.
02
We turn it into short tasks
Your data becomes quick judgments vetted reviewers make on their phones, with the context right beside them.
03
The crowd judges, consensus decides
Several independent reviewers per item, hidden attention checks, trust scores, and an expert for anything disputed.
04
Follow along and download
Watch progress in your client portal, then download CSV or Excel with your own IDs, plus a PDF or PowerPoint report.
No sample metrics here
You won't find invented percentages on this page. Every number we give you is counted from your own data, with the agreement behind it.
Tell us what you need checked. We'll design the review.
Share a few example conversations or items and what you want to learn from them. We'll come back with a proposed rubric, how we'd run the review, timelines and pricing.