Inside the platform

How the work actually gets done.

AfriEval joins two halves that are usually kept apart: the people who judge whether a model is right, and the machinery that runs, records and exports the evaluation. This page walks through both.

Reusable workflow builder
Build a task once from columns, groups, sections and side by side comparisons, then run it across many projects. Workflows are the design; projects are the live executions.
Two layer quality
Contributors certify in the language they work in, then pass an exam written by the client for that specific project. Senior reviewers adjudicate answers row by row.
Corpora at scale
Import, index and search document collections in the hundreds of thousands. Import, keyword indexing and embeddings each run as independent workers that resume where they stopped.
Traceable evaluation runs
Every model run records what it searched, what it read and what it answered, so a result can be reopened and explained rather than taken on trust.
How a project runs

Six steps, start to export

The same path whether a project is fifty rows for a research paper or a hundred thousand for a production model.

01
Design the workflow
Choose the question components, the guidance contributors see, and the fields that get written back to the dataset.
02
Bring the data
Upload a spreadsheet or document collection. Columns are mapped to the workflow, and rows are grouped into tasks.
03
Set the quality bar
Pick required certifications, add a project exam, set how many contributors see each item, and how disagreements escalate.
04
Staff the work
Open the project to certified contributors, or keep it to a private pool you invite by email.
05
Review and adjudicate
Senior reviewers read answers row by row, resolve conflicts, and sign off on what ships.
06
Export
Take results out as Excel or JSON lines, formatted for training, analysis or a public benchmark submission.
Quality

Two gates, then a human sign off.

Nobody starts on paid work by accident. Certification proves the language, the client exam proves the task, and senior review is where the final call is made.

Layer one
Language certification
Run by AfriEval and tied to a specific language. A contributor certified in Kiswahili is not automatically cleared for Yorùbá.
Layer two
Client qualification exam
Written by the client for their own project, so the bar reflects the guidelines that project will be judged against.
Then
Row by row senior review
Senior reviewers read each answer next to the item it belongs to, resolve disagreements, and export only what has been signed off.
Model evaluation

Runs you can reopen and explain

Questions are asked in an African language and answered from your own documents. Every search round and every tool call is recorded, so a disputed answer can be traced back to what the model actually read.

Your own model keys
Connect your accounts across the major providers. Keys are encrypted at rest, and the choice of model belongs to the run, not the key.
Search over your documents
Questions are answered from your own collection. Keyword search runs today; ranked fusion with semantic search is built and switched on per run.
Answers that cite sources
Every answer names the documents it came from, and anything unconfirmed is flagged rather than smoothed over.
Benchmark ready output
Exports are checked against the benchmark's own validator before you submit, so formatting problems surface here rather than there.
Built for scale
100,000+
documents imported in a single collection
Resumable
import, indexing and embeddings each continue where they stopped
Validated
exports checked against the benchmark's own validator before submission

Who it is for

Contributors
Certify in your language, build a quality record, and work on projects that match your expertise.
Researchers
Run reproducible evaluations in African languages and publish results with the full trace attached.
Teams building AI
Collect, annotate, evaluate and benchmark in one place, with permissions and an audit trail throughout.

See it on your own data.

Bring a question set and a document collection, and we will run an evaluation end to end with the full trace attached.