Annotator quality
This report briefly describes the annotation quality-control system used by Rapidata to collect human feedback at scale.
1. How annotators encounter tasks
Rapidata uses consumer application and advertising distribution to expose short annotation tasks to a large international population. Users may see a compact evaluation interface embedded in consumer apps, including apps such as Duolingo. Part of the exposed users choose to interact with the task. The qualification process below determines which judgements are eligible to count.2. Acquisition to graduation
New annotators begin with a set of validation tasks. Their benchmark judgements count only after they have answered enough validation tasks correctly to qualify for the global audience (or any custom audience specific to some benchmarks). Qualification is therefore based on demonstrated task performance, not simply on participation.3. Validation tasks
A validation task has a known answer and is presented through the same interface as a benchmark task. These tasks are designed to capture whether an annotator understands the evaluation criterion, pays attention to the content, and can apply the requested judgement consistently rather than answering randomly or with low effort.Validation tasks are used during initial qualification and continue to appear throughout an annotator's activity. This allows reliability to be reassessed over time and helps identify changes in attention, consistency or task understanding.
4. Global audience qualification
The global audience is Benchmark.ai’s general-purpose pool of annotators who have qualified and remain above the required reliability threshold by answering validation tasks. Public benchmarks use this audience by default unless the benchmark protocol specifies a curated or custom audience.5. Continuous reliability measurement
Each annotator carries a score that is updated continuously based on their validation-task performance and other quality signals. For every accepted workflow task, Rapidata distributes approximately two validation tasks, allowing reliability to be measured throughout the annotator's activity rather than only during initial qualification. The score reflects whether the annotator's judgements can be trusted and remains eligible to contribute to benchmark results.6. Curated and Custom Audiences
Public benchmarks normally use the global audience. A curated audience selects qualified annotators using validation questions relevant to the evaluation, such as prompt alignment, coherence etc… A custom audience may additionally use project-specific validation questions. Whenever a curated or custom audience is used, the benchmark protocol should identify the applicable selection criteria.7. Multiple independent judgements
Each pairwise match is evaluated by several distinct annotators. By default, five independent judgements are collected for each match. Model identities are hidden from annotators, and the left-right order of the outputs is randomised for every judgement. Collecting several independent judgements reduces the influence of any single inconsistent response.8. Auditability
Rapidata retains the information needed to audit workflow outcomes at the datapoint level. For each evaluated item, the system can track which votes were included, which annotators cast those votes, the raw vote distribution, the reliability-weighted vote distribution, and the resulting uncertainty estimate. This makes it possible to inspect the empirical basis of a benchmark score or customer result rather than treating it as an opaque aggregate.9. Annotator demographics
Self-reported demographics vary per round because the audience is resampled. A report describing overall annotator demographics over the past few months will be released soon.The task, as shown to an annotator
Figure 2. Task interface: live previews of an image and a video task. Model identities are hidden and left/right order is randomized per judgement.
We use crowd intelligence to deliver real, diverse human feedback at scale for model evaluation and post-training.
Rapidata, benchmark.ai © 2026