The annotation queue is growing faster than the model team can clear it. Engineers are spending Fridays reviewing labels, domain experts are being pulled into repetitive decisions, and a supplier has offered to process the backlog at an attractive unit price.
The quote may solve the visible capacity problem. It does not answer the more important question: which decisions can safely leave the company?
Annotation is not simply repetitive data preparation. Labels encode assumptions about customers, risk, product behaviour, and acceptable error. A provider can supply labour and process discipline, but the company remains responsible for what the resulting model learns.
The real decision is what to retain, not where to send the work
The common framing—internal team versus external provider—is too coarse. A better operating question is:
Which parts of data production contain proprietary judgement, sensitive access, or regulatory accountability, and which parts can be delivered by a measurable external workforce?
That distinction matters because annotation programmes combine several different activities:
- defining the ontology or label taxonomy;
- writing task instructions and examples;
- preparing and minimising source data;
- completing first-pass annotations;
- reviewing difficult or high-risk cases;
- adjudicating disagreements;
- monitoring errors and model effects;
- documenting lineage, access, and changes.
Few companies should outsource that entire chain. First-pass labelling and structured review may be suitable for a provider. Taxonomy ownership, final adjudication, risk acceptance, and model evaluation usually need an accountable internal owner.
Growing AI adoption does not change that requirement. Statistics Canada reported that 12.2% of Canadian firms used AI to produce goods or deliver services in 2025, but its firm-level analysis found that the apparent productivity advantage became statistically insignificant after accounting for prior productivity and complementary capabilities such as R&D, cloud systems, analytics, and training (Statistics Canada). External annotation capacity only creates value when it is part of a functioning delivery system.
When external annotation is a sensible capacity choice
Outsourcing tends to fit when the task has clear boundaries, repeatable instructions, sufficient volume, and measurable acceptance criteria.
Typical signals include:
- a backlog is delaying model training or evaluation;
- demand changes enough that permanent internal headcount would sit idle between batches;
- the work spans languages, time zones, or modalities that the current team cannot cover;
- engineers and domain specialists are spending significant time on routine labels;
- data can be redacted, pseudonymised, synthesised, or segregated before access;
- the company can maintain an internal gold set and review function.
The case is weaker when ambiguity dominates volume. If every twentieth item requires product context, clinical judgement, legal interpretation, or a discussion with engineering, a distant production queue may create more coordination than capacity.
Keep the work internal, or use a tightly controlled hybrid model, when the taxonomy is defensible product knowledge, the data cannot be safely minimised, volumes are low, or the cost of a wrong label is disproportionate. Safety decisions and rare-event classification deserve particular caution because aggregate accuracy can conceal serious class-level errors.
Outsourcing is also premature when nobody internally owns the instructions. A supplier cannot stabilise a taxonomy that changes informally from one stakeholder to another.
Build the internal control layer before buying volume
A provider relationship needs a small but credible internal control layer. Depending on the use case, one person may cover several responsibilities, but each responsibility should be explicit.
Roles that should have named owners
Data or ML product owner: connects annotation priorities to model and product outcomes. This person decides which data is worth labelling and whether throughput is improving delivery.
Data operations lead: owns workflow design, batch planning, instructions, quality reporting, and provider cadence.
Domain adjudicator: resolves ambiguous examples and approves changes to the taxonomy. This may be a clinician, lawyer, safety specialist, linguist, or experienced product operator.
Privacy or security owner: approves access patterns, processing locations, retention, subprocessors, and incident procedures.
ML evaluation owner: checks whether accepted labels improve class-level model performance rather than simply passing a production metric.
Without this layer, the vendor manager becomes a message forwarder. Instructions drift, edge cases accumulate, and the company struggles to explain why a label was applied.
The control layer does not need to be large. It does need enough authority and time to maintain gold examples, issue versioned guidance, inspect performance, and stop a batch when systematic errors appear.
Define quality before asking for a price
A promise of high accuracy is not a quality plan. Accuracy against which reference, across which classes, and with what treatment of ambiguous items?
A 2024 review of 591 publications introducing text datasets found that 30% showed subpar quality-management effort. The researchers also identified recurring problems in how inter-annotator agreement and annotation errors were calculated or interpreted (Computational Linguistics). The lesson for buyers is straightforward: workforce size and a headline percentage do not establish dataset quality.
Use a task-specific scorecard containing several measures:
- Gold-set agreement: performance against examples approved by your domain owner.
- Class-level error: false positives and false negatives for each important category.
- Severity-weighted error: separate harmless mistakes from errors that could affect safety, compliance, or customer outcomes.
- Reviewer overturn rate: how often accepted labels are changed during internal review.
- Inter-annotator agreement: useful for detecting ambiguity, but not a substitute for correctness.
- Escalation rate: the share of items annotators correctly identify as uncertain.
- Rework time: internal and external effort required before a batch is accepted.
- Instruction drift: changes in performance after guidance, workforce, or taxonomy updates.
Include ordinary, rare, ambiguous, and adversarial cases in the test set. A provider that performs well only on the obvious majority class will look efficient until the model reaches production.
Treat privacy and chain of custody as operating requirements
Sending data to an external workforce expands the number of people, systems, and organisations involved in processing it. The relevant questions are concrete: what can each annotator see, from where, for how long, and through which system?
Canadian private-sector guidance does not prohibit cross-border processing, but it states that the organisation transferring personal information remains accountable and should use contractual or other measures to provide comparable protection (Office of the Privacy Commissioner of Canada). Provincial and sector-specific requirements may add further constraints.
For European businesses, annotation can form part of a more explicit governance obligation. Article 10 of the EU AI Act includes annotation, labelling, cleaning, bias examination, and mitigation within the data-governance practices required for high-risk AI systems (EU AI Act).
Where personal data is involved, GDPR Article 28 requires controllers to use processors offering sufficient guarantees and to put a processing contract in place. That contract must address the processing purpose and duration, data types, data-subject categories, obligations, and conditions for subprocessors (GDPR, Article 28).
Before a pilot begins, document:
- processing and storage locations;
- every authorised subprocessor;
- role-based access and authentication controls;
- restrictions on local downloads, screenshots, and personal devices;
- logging and review arrangements;
- retention periods for source data, labels, backups, and derived files;
- incident notification and investigation duties;
- data return, deletion evidence, and transition support at exit.
A security questionnaire is only the starting point. Test the actual operating path with the same care as the contractual wording.
Run a bounded pilot before committing the backlog
A useful pilot is reversible and difficult enough to reveal operating weaknesses. Select one modality and use case, minimise the data, include known edge cases, and compare the provider against an internal baseline.
Four weeks may be enough to observe training, production, review, and correction cycles, but the right duration depends on task volume and complexity. Do not expand simply because the provider completed the first batch quickly.
Worked example: retail image annotation
Consider a European computer-vision team with 10,000 de-identified retail-shelf images awaiting classification each month. Its taxonomy includes common products and a small number of rare shelf conditions that matter to model performance.
The team retains taxonomy design, gold-set ownership, rare-class adjudication, and weekly error review. A provider receives a fixed pilot batch containing ordinary images, rare cases, blurred images, and deliberate ambiguities.
Assume the following illustrative internal baseline—not an industry benchmark:
- 500 labelling hours at €45 fully loaded per hour: €22,500;
- 80 hours of senior review at €65 per hour: €5,200;
- 9,000 labels accepted without further rework;
- fully loaded cost per accepted label: €27,700 ÷ 9,000 = €3.08.
Now suppose the vendor quotes €1.40 per submitted image. The comparison should not stop there. Add internal review time, provider-management time, rejected labels, rework charges, tooling, secure-environment costs, and transition work.
If the external pilot costs €14,000 for production, €7,800 for internal review and management, and €1,500 for rework, while delivering 8,700 accepted labels, the fully loaded cost is approximately €2.68 per accepted label. That may justify further testing, but only if rare-class errors, documentation, and access controls meet the agreed thresholds.
The decision branches are then clear:
- Scale the provider if difficult-case quality, control evidence, throughput, and economics all pass.
- Use a hybrid pod if routine labels pass but rare or sensitive cases need internal handling.
- Hire internally if coordination and rework erase the labour saving, or if product judgement dominates the task.
Use this checklist before approving a supplier
Work design
Quality and economics
Privacy, security, and continuity
A supplier that cannot answer these questions during a bounded pilot is unlikely to become easier to govern at ten times the volume.
Make the next decision: supplier, internal owner, or both
The right first move is not always issuing an annotation tender. If your taxonomy is unstable, quality disputes remain unresolved, or no one owns provider governance, hiring the internal control role should come first.
A senior data operations lead can define the workflow and scorecard. An ML product owner can connect labels to model outcomes. A domain specialist can establish adjudication policy. A privacy or security specialist can shape an access model that a provider can actually follow.
Deeptal connects companies with senior European specialists for roles such as these. Shortlists are human-reviewed, there are no recruitment fees, and Deeptal can handle contracts, payroll, compliance administration, and consolidated monthly billing. For a qualified brief, an initial shortlist is typically prepared within two business days; selected specialists typically start in 7–14 days, subject to fit, availability, interviews, and terms.
Before buying more annotation capacity, decide who inside the company will own the system around it. If that role is missing, start the hiring brief there.



