A product team is six weeks from release. Its model performs well in the aggregate, but misroutes German billing requests and struggles with ambiguous French tickets. Engineering wants more labelled data. Procurement has three supplier proposals, each promising high accuracy, experienced annotators and layered quality assurance.
The proposals look comparable. They are not.
The real difference will emerge in how each supplier handles uncertain cases, guideline changes, linguistic nuance and evidence. A provider can complete every item on time while producing a dataset that weakens the model, obscures risk and creates expensive rework.
That is why supplier selection should not begin with price per label. It should begin with a production-relevant task, a buyer-owned definition of quality and a paid pilot designed to expose failure modes.
A provider cannot define quality for you
“High-quality data” is not a useful acceptance criterion. Quality depends on the model’s intended use, the consequences of error and the distribution it will encounter after release.
For a retail image classifier, quality might mean consistent product categorisation across lighting conditions and camera types. For a medical workflow, it may require specialist judgement, documented uncertainty and escalation. For a multilingual support model, average label accuracy matters less if errors concentrate in a commercially important language or issue type.
A credible specification should therefore cover at least four dimensions:
- Correctness: Do labels or evaluations agree with a defensible reference standard?
- Consistency: Do different annotators apply the same rule to comparable examples?
- Representativeness: Does the dataset reflect languages, groups, edge cases and operating conditions that matter in deployment?
- Traceability: Can the team reconstruct where an item came from, which guideline version governed it and how disputes were resolved?
NIST connects AI failure modes to data quality and representativeness and recommends documenting the assumptions behind data selection, preparation and analysis (NIST AI RMF Playbook). That makes slice-level analysis essential. One overall accuracy figure can conceal poor results on rare classes, noisy images, regional language or high-impact cases.
Claims about trained annotators, peer review, domain-expert QA and feedback loops are reasonable screening inputs. They are not proof. Buyers need to see how those controls operate on their own task and data.
Decide whether outsourcing is the right operating model
A specialist data operation is most useful when three conditions are present:
- The use case and human task can be defined with reasonable precision.
- The work recurs at enough volume to justify dedicated operations.
- Internal engineers or domain experts are becoming a delivery bottleneck.
Examples include document classification, search relevance judgement, safety evaluation, retrieval assessment, red-team review and recurring human-feedback workflows.
Outsourcing is less suitable when the taxonomy changes every week, the task depends on scarce internal knowledge, or the labelling decisions themselves represent core intellectual property. It may also be disproportionate for a small, one-off dataset. In those cases, a compact internal expert team should establish the ontology, gold set and escalation policy before external capacity is added.
The question is not simply whether a supplier can label faster. It is whether the work is stable enough to transfer without losing the reasoning behind it.
Research on “data cascades” in high-stakes AI found that upstream data problems can compound across later stages of development. The study was qualitative rather than a universal market estimate, but it illustrates why task design should be treated as engineering work rather than low-cost administration (Google Research).
Specify the task before requesting prices
A weak request says: “We need 100,000 support tickets categorised.”
A decision-ready brief explains:
- the model or evaluation use case;
- the unit of work;
- the taxonomy or scoring rubric;
- permitted and prohibited data use;
- known sources of ambiguity;
- languages and domain expertise required;
- the cost of false positives and false negatives;
- expected volume and throughput;
- acceptance metrics and reporting slices;
- escalation rules for uncertain cases;
- required outputs, metadata and audit records.
The buyer should own the intended-use statement, core ontology and final decision rights. A supplier may improve instructions, but it should not quietly redefine the product policy embedded in the labels.
Build a small gold set before approaching providers. Internal domain experts should label it independently, record their rationales and adjudicate disagreements. Include easy items, boundary cases and examples where competent people could disagree.
A gold set is not an eternal truth. It is a controlled reference point. When product policy changes, the set and its guideline version should change together.
ISO/IEC 5259-1:2024 provides a neutral vocabulary for discussing AI and machine-learning data quality across the lifecycle. It can help procurement, engineering and suppliers use consistent terms in an RFP or service schedule (ISO).
Inspect the provider’s quality system
Do not ask only, “What accuracy can you achieve?” Ask the provider to show how its operation detects, explains and corrects errors.
The evidence pack should address:
- Annotator selection: How are language, domain and task skills assessed?
- Calibration: What happens before production and after disagreement rates increase?
- Agreement measurement: Which metric is used, and why is it suitable for this task?
- Adjudication: Who decides difficult cases, with what expertise and response time?
- Sampling: How are completed items selected for review? Are high-risk slices oversampled?
- Error taxonomy: Are mistakes classified by cause rather than counted only as failures?
- Change control: Can the provider identify work affected by a revised instruction?
- Performance analysis: Are results broken down by class, language, annotator cohort and edge-case type?
- Corrective action: Who pays for rework, and how is the root cause addressed?
Inter-annotator agreement is useful, but it is not sufficient. Two annotators can consistently apply the same flawed instruction. Agreement must be considered alongside expert adjudication, gold-set performance and model impact.
Lineage matters as well. NIST recommends documenting data provenance, including origin, transformations, labels, dependencies, constraints and relevant metadata (NIST AI RMF Playbook). Make those records exportable. If the relationship ends, the buyer should retain enough information to reproduce the dataset and understand historical decisions.
Use a paid pilot as an evidence gate
A short unpaid sample often encourages suppliers to assign unusually strong staff to a tiny batch. A paid, time-boxed pilot is more revealing because it tests the operating model under realistic conditions.
Use six gates:
Scroll horizontally to see all columns.
| Gate | Buyer action | Evidence required to proceed |
|---|---|---|
| 1. Define | Select one production-relevant task and issue a versioned guide. | Intended use, edge-case policy, output format and acceptance metric. |
| 2. Build the reference | Create a representative gold set with internal experts. | Gold labels, rationales, known ambiguities and slice definitions. |
| 3. Run blind | Give each supplier the same blinded sample with limited intervention. | Completion rate, turnaround, gold-set results and disagreement data. |
| 4. Test change | Introduce a controlled guideline revision or difficult batch. | Calibration record, affected-item analysis and updated version history. |
| 5. Review data flow | Trace where the data enters, who can access it and where it leaves. | Contract terms, access controls, subprocessors, retention and deletion process. |
| 6. Decide | Compare quality, operating evidence, risk and total cost. | Negotiated service levels, correction rights, ownership and exit plan. |
Do not coach suppliers after every error. That tests the buyer’s supervision, not the supplier’s system. Allow normal clarification, but record each question and assess whether the eventual answer becomes part of the controlled guidance.
The pilot should also resemble future production. If the live dataset contains six languages and irregular scans, a clean English-only sample proves little.
Worked example: multilingual support routing
Consider a European software company preparing a model that routes incoming support tickets. It pilots two suppliers on 2,000 anonymised tickets covering four languages, eight issue categories and a mixture of clear and ambiguous requests.
The internal team assigns weights before seeing results:
Scroll horizontally to see all columns.
| Criterion | Weight | Provider A | Provider B |
|---|---|---|---|
| Gold-set correctness | 40 | 34 | 31 |
| Performance on critical slices | 20 | 11 | 18 |
| QA traceability and change control | 15 | 10 | 14 |
| Security and data handling | 15 | 14 | 13 |
| Commercial fit | 10 | 9 | 7 |
| Total | 100 | 78 | 83 |
Provider A has better overall agreement and a lower unit price. Provider B performs better on German billing and French account-access cases, records adjudication rationales consistently and identifies which earlier items are affected when the team changes a category definition.
If those categories carry high support cost or customer risk, Provider B is the stronger choice despite its higher price and slightly lower aggregate score.
The weighting should change with the use case. A low-risk catalogue task may place more weight on throughput and cost. A safety evaluation or regulated workflow should place more weight on specialist competence, traceability and controlled escalation.
The point of the matrix is not mathematical certainty. It prevents a single attractive metric from deciding the purchase.
Contract for evidence, not aspiration
Once a supplier passes the pilot, translate the tested operating model into the contract and service schedule. Avoid broad promises such as “industry-leading accuracy” without a defined sample, metric, threshold and remedy.
Where personal data are involved, this is also a processor-selection decision. GDPR Article 28 requires controllers to use processors that provide sufficient technical and organisational guarantees, with binding terms covering the processing and controls around subprocessors (EUR-Lex). Security and privacy teams should review the actual data flow rather than relying solely on a questionnaire.
Use this contracting checklist:
For high-risk AI systems in Europe, dataset governance can also become a product obligation rather than an internal preference. Article 10 of the EU AI Act sets data and data-governance requirements for relevant training, validation and testing datasets (EUR-Lex). Legal applicability depends on the system and its classification, so obtain qualified advice rather than treating a supplier certificate as a substitute.
Make the next decision about ownership
Before opening a supplier process, assign one internal owner for the quality system. That person should control the task definition, acceptance criteria, guideline versions, pilot design and final adjudication rights. Procurement can negotiate price and terms, but it cannot decide what a correct label means.
If your task is already stable and repeatable, proceed to a paid supplier pilot. If it is not, the immediate need may be a senior applied AI or ML specialist who can build the evaluation harness, data checks, privacy controls and human-feedback workflow before volume work begins.
Deeptal is not presented as a managed annotation BPO. It connects companies with senior European specialists who can own the technical layer around a data supplier. There are no recruitment fees, and Deeptal can handle contracts, payroll, compliance administration and consolidated monthly billing. For a qualified brief, a human-reviewed initial shortlist is typically prepared within two business days; selected specialists typically start within 7–14 days, subject to fit, availability, interviews and terms.
The useful next step is to write down the production task, the unresolved quality risk and who currently owns acceptance. If that reveals an engineering ownership gap, start a Deeptal hiring brief around that gap rather than buying annotation volume prematurely.



