Trust

Responsible AI starts with responsible data.

Before a model can be responsible, the data behind it must be lawful, traceable, culturally grounded and human-governed.

CIHAN AI DATA is an African AI-data, annotation, human-evaluation and benchmark operation. Our risk profile sits upstream of model behaviour: how data is sourced, who interprets it, how disagreement is resolved and how results are described. These commitments govern the way we run commercial projects and research alike.

01Provenance before volume

Every dataset carries a record of where it came from, how it was collected and what permissions apply. We would rather deliver a smaller corpus with defensible provenance than a large one that cannot be explained.

02Consent and lawful sourcing

We do not run covert or unlawful collection. Each project has written sourcing rules, and where consent is the appropriate basis we obtain and record it for the specific purpose described to participants.

03Cultural context by design

African language, register, dialect and code-switching work is assigned to locally grounded analysts using documented guidelines, so cultural meaning is interpreted rather than approximated from outside the context.

04Human dignity and fair work

Contributors are managed network talent, not an anonymous crowd. Voluntary research participation is separated from paid commercial work, and commercial production tasks are compensated under project terms.

05Human evaluation and reviewer escalation

Calibrated contributors work in reviewer tiers. Disagreements are surfaced rather than averaged away, and contested items escalate to senior review and adjudication with the rationale recorded.

06Benchmark integrity

Holdouts are protected, contamination controls are applied, gold labels are access-restricted, benchmark versions are tracked, and every score is reported against the benchmark that produced it.

07Privacy and security

We apply data minimisation, role-based access and project segregation. Sensitive material is handled only under explicit project terms and restricted to the people who need it for delivery.

08Contestability and correction

Clients, contributors and affected individuals can raise a data, quality, bias, privacy or benchmark concern and get a documented response, including correction where the concern is upheld.

Common questions

What does responsible data mean at CIHAN AI DATA?
It means every dataset we touch has a documented origin, a lawful basis for collection, culturally grounded annotation guidelines and named human accountability for quality — before any volume target is discussed.
How are African language and dialect tasks handled?
Register, dialect and code-switching tasks are assigned to locally grounded analysts working from documented guidelines, with calibration and adjudication rather than a single unreviewed pass.
How is benchmark integrity protected?
Holdout items and gold labels are access-restricted and versioned, contamination controls are applied to benchmark design, and results are reported as benchmark-specific rather than as universal model rankings.
Are contributors paid?
Voluntary research participation and paid commercial project work are clearly distinguished. Commercial production work is compensated under the applicable project and contributor terms.