Crowdworkers for AI Evaluations: Platforms, Costs, Expert Review, and a Privacy-Safe MyVault Playbook


An AI system can produce a fluent answer in seconds; establishing whether that answer is correct, safe, or privacy-preserving typically requires people who work under very different constraints. Reviewers vary by who recruited them, what they are paid, what credentials they hold, whether they work on a managed contract or an open task queue, and what technical controls prevent them from copying or aggregating sensitive information. That mismatch — instant model outputs versus slower, institutionally governed human judgment — is the practical tension MyVault must resolve before any external human review occurs.
For clarity in this article: “AI evaluations” are structured checks of model outputs or retrievals to measure correctness, safety, or user experience; “annotation/labeling” is the process of assigning categorical tags or extracted fields to items; “preference ranking” (pairwise judgments) asks reviewers to choose which of two outputs is better on defined criteria; and “human review” covers senior adjudication, rubric design, or expert analysis that resolves disagreement or inspects edge cases. These activities overlap but demand different skills, sampling rules, and access controls: a paid, recruited participant can rate usability, while a credentialed adjudicator must resolve legal or medical interpretations. The governance themes in NIST’s Generative AI Profile emphasize program structure, monitoring, and assigned roles for evaluations but do not set legal requirements; they are practical guidance to adapt with accountable owners and retained records (NIST AI 600‑1).
MyVault’s core operating rule (pilot proposal): outsource privacy‑preserving evaluation, not possession of the vault. As a default stance for the pilot described here, external reviewers should normally receive synthetic cases, strongly de‑identified excerpts, or derived model outputs. They should not receive raw family documents, account credentials, full chat histories, identity mappings, or unrestricted product access. De‑identification may reduce, but does not automatically eliminate, disclosure risk; whether a transformed packet is still personal data can depend on context, auxiliary information, and applicable law. NIST’s de‑identification guidance describes protected sharing models and cautions that risk assessments remain necessary (NIST SP 800‑188). If MyVault operates in jurisdictions such as the EU/EEA, data protection rules (e.g., the GDPR) may apply to evaluation workflows; controller‑processor roles, transfer conditions, and DPIA triggers are fact‑specific and should be reviewed with counsel (GDPR text). Nothing in this article is legal advice.
Public pricing evidence for external human work is narrow. The participant‑panel provider Prolific describes a buyer charge equal to the participant reward plus a platform fee (their published example lists a 42.8% platform fee for corporate buyers); Clickworker’s survey product describes a buyer‑set participant fee (US$0.25 minimum) plus a 40% service fee (Prolific pricing; Clickworker survey respondents). Those pages are useful as planning anchors for lower‑sensitivity research but do not generalize to expert adjudication, managed red teaming, or secure‑workspace programs, which are typically quote‑only.
Readers should expect a practical, role‑specific task mapping and a procurement checklist: which tasks are safe for public or recruited crowds, which require named experts, and which should remain inside MyVault or in a documented managed‑service enclave. Subsequent sections describe four workforce models, a defensible cost framework with public price anchors where they exist, quality controls to treat human labels as measurements rather than opinions, and a 90‑day MyVault pilot that enforces the “no possession” privacy gate as a proposal to test in practice.
The Four Human‑Work Models Behind AI Evaluation
The market often grouped under “crowdwork for AI” contains at least four distinct operational models: self‑serve crowds, recruited participants, screened expert networks, and managed evaluation services. Procurement and risk decisions should be based on the delivery model and contract controls, not on a headline “crowd size.” Predictable differences — who recruits and trains reviewers, who owns the rubric and adjudication, whether buyer data leaves a controlled workspace, and whether pricing is posted or quote‑only — tend to shape outcomes far more than nominal contributor counts.
| Model | Best‑fit examples | Main advantage | Main risk | MyVault pilot proposal |
|---|---|---|---|---|
| Self‑serve crowd | High‑volume, low‑sensitivity tasks (relevance, basic labels) | Fast scale and low setup friction | Weaker device/location controls; public‑crowd confidentiality risk | Only synthetic or strongly de‑identified packets; never raw vault data |
| Recruited participants | Representative UX, comprehension, preference studies | Improved sampling/provenance vs open queue | Not inherently an annotation/secure platform; external study tools may handle data | Use for sanitized UX and preference tests, with MyVault‑controlled task app |
| Screened expert network | High‑reasoning, domain‑specific adjudication or rubric design | Higher domain fit for edge cases and specialist judgment | Credential depth and availability may be opaque; expertise ≠ privacy | Use for synthetic high‑risk cases, rubric authoring, and named adjudicators |
| Managed evaluation service | Large programs, secure‑facility work, regulated review, red teaming | Dedicated staff, workflow customization, contractual accountability | Quote‑based pricing, longer procurement, dependence on provider statements | Consider if residual‑sensitive review is unavoidable after strict diligence |
Compact public pricing and worker‑pay evidence are limited and typically apply to narrow products rather than full evaluation programs. For example, Prolific publicly posts a corporate fee structure that adds a platform fee to participant rewards, and Clickworker’s survey product documents buyer‑paid participant minimums plus a service fee (Prolific pricing; Clickworker survey respondents). Those anchors can help scope sanitized preference or comprehension work; they are not guarantees for secure‑facility review, expert adjudication, or red‑team engagements.
Hybrid providers and practical implications
Many vendors span more than one model in their product mix. Their public materials often describe combinations of self‑serve queues, managed projects, expert sourcing, and configurable workflows. Treat such statements as vendor claims to verify through proof points and contracts rather than as settled facts. Practically, this means procurement should map the exact engagement mode — who sees the data, what workspace is used, who adjudicates, and how pay and retention are handled — rather than accepting a vendor’s single label.
For MyVault, the pilot position proposed here is conservative: route low‑risk work to recruited or public self‑serve panels using only D0/D1 (synthetic or strongly de‑identified) packets, and reserve experts or managed services for controlled staging‑only evaluations after contractual and technical evidence of data minimization, short retention, deletion attestations, and workspace controls are obtained and tested.

What Crowdworkers and Experts Actually Do in an AI Eval Flywheel
AI evaluation is a bundle of distinct human tasks — each with different skill, privacy, and audit implications. Below are the common task types, a concise statement of the work involved, and which reviewer cohort is typically appropriate in a privacy‑first pilot like MyVault’s.
Data annotation and labeling. Workers apply stable, bounded labels (document type, entity spans, OCR verification) against a rubric. These tasks reward throughput, clear instructions, and sentinel (“gold”) controls. Self‑serve crowds and managed annotation pools are commonly used for scale and multilingual coverage, while expert networks and managed services are often brought in for complex or regulated document types. Because annotation platforms and vendor practices vary, buyers should insist on calibration, gold‑task rotation, and adjudication as explicit line items and should test the workspace controls during a small pilot run.
Factuality checks and grading. Reviewers verify factual claims against sources or mark hallucination severity. Because factual checks sometimes require domain knowledge or source‑quality judgment, a two‑tier approach is pragmatic: general reviewers for surface factuality and provenance checks on sanitized outputs, and credentialed experts for domain‑sensitive facts (medicine, law, finance) or novel evidence conflicts. OpenAI’s public discussion of instruction‑following highlights demonstrations, rankings, screened labelers, and the limits of average rater preference for truth — useful reminders to separate preference from verifiable correctness (OpenAI: instruction‑following).
Search and retrieval relevance; RAG grounding. Relevance tasks ask whether a retrieved excerpt answers a query; retrieval‑augmented generation (RAG) grounding tasks assess whether an answer cites or matches supporting passages. Recruited participants or general annotators can be suitable for representative relevance and UX judgments when packets are synthetic or strongly de‑identified. If grounding involves residual‑sensitive real excerpts or high‑consequence misattribution, restrict access to a small, named expert cohort in a controlled workspace with short retention and no secondary use, and perform a privacy assessment first (see de‑identification and protected sharing models in NIST SP 800‑188).
Pairwise preference and ranking. Preference tasks present two outputs and ask which is better. Pairwise work is efficient for perceived quality or helpfulness, but it can reward style over factual grounding. Design should separate preference from correctness and include reason codes or brief rationales. OpenAI’s public materials discuss demonstrations and pairwise rankings as part of training/evaluation practices, which buyers can adapt with caution for evaluation purposes (OpenAI).
Rubric authoring and gold‑set creation. Writing clear label definitions, boundary cases, and adjudicated gold items is specialized work. Managed vendors or specialist networks are frequently used for rubric design and the creation of rotating sentinel items to detect drift; these activities are often billed as discrete statement‑of‑work items and should be performed by named reviewers with documented qualifications.
Adversarial red teaming. Red teams design attacks (prompt injection, privacy leaks, false authority) and attempt to elicit unsafe outputs. Public research from Anthropic illustrates that large‑scale red‑team work can surface disagreements and edge behaviors that require iterative oversight and re‑rating, not a one‑time sweep (Anthropic red‑team research). For a privacy‑sensitive product like MyVault, treat red teaming as an expert or managed‑service activity in staging, not a general crowd task, and gate any real‑data exposure behind explicit technical and contractual controls.
Multilingual and cultural review. Native speakers assess idiomatic naturalness, cultural sensitivity, and translation fidelity. Self‑serve crowds can provide breadth for sanitized items, whereas expert reviewers and managed providers are better suited where jurisdictional, legal, or cultural nuance materially changes the interpretation of content. Use calibration and adjudication to track slice‑level performance by language.
Expert adjudication. Senior reviewers resolve disagreements, overturn incorrect consensus, and update rubrics. Adjudicators record rationale, tag failure modes, and trigger rework or regression tests. In the MyVault pilot proposal, adjudication is a named role that remains under MyVault’s authority even when some external judgments are used.
The operational flywheel — concrete steps
A practical evaluation program is iterative: observe failures, sample, get independent judgments, adjudicate, repair, convert accepted failures into tests, and repeat. The NIST Generative AI Profile encourages assigning oversight roles, retaining evaluation records, and monitoring performance as part of a broader risk‑management posture (NIST AI 600‑1).
- Observe a failure or anomaly in production or a staging run.
- Sample representative and risk‑targeted cases (a random monitoring stream + a risk‑directed stream).
- Collect independent judgments (redundant crowd or recruited reviewers; experts for edge cases).
- Escalate disagreements to named adjudicators who record rationale and prescribe rubric changes.
- Implement repairs (model, retrieval, rubric, or pipeline), add regression tests from adjudicated failures, and re‑test a frozen holdout.
Each loop should produce audit metadata: rubric version, reviewer cohort, gold accuracy, inter‑reviewer consistency by slice, overturn rate, and a retention record for regression tests. In the MyVault pilot proposal, calibration work is paid, qualification rounds precede production, and a frozen holdout is re‑tested after remediation so improvements are distinguishable from overfitting to visible gold sets.
In practice, these roles are complementary: crowds provide breadth and diversity for large‑scale preference and relevance checks; recruited participants add demographic or UX representativeness; expert networks and managed services supply adjudication, rubric engineering, red teaming, and any residual‑sensitive review under strict contractual and technical controls. MyVault should retain ownership of identity linkage, final authority on privacy incidents, and the frozen regression suite.
For implementation patterns, see this guide to human‑in‑the‑loop workflow design, where asynchronous questions and bounded approvals preserve accountable human control.
Provider Landscape: Market Signals and Diligence Priorities
How to read this market: separate providers by operating model, not by headline crowd size. Ask whether a provider operates a self‑serve task queue, a recruited participant panel, a screened expert network, or a managed service that owns recruitment, training, adjudication, and facilities. Public pages and role ads are useful signals, but they are not proof; treat provider claims as statements to verify through a small pilot and through contract language.
Three different price categories: (1) public buyer list prices (rare and product‑specific), (2) worker pay signals (advertised ranges to contributors), and (3) negotiated enterprise quotes. The public formulas for Prolific and Clickworker surveys are usable planning inputs for sanitized participant tasks; most managed AI‑evaluation work is quote‑only and requires line‑item RFPs to reveal setup, security, adjudication, and retention charges (Prolific pricing; Clickworker survey respondents).
Primary routing axis: data sensitivity. For MyVault — with family and personal information — the default pilot proposal is to avoid sending raw personal or control‑plane data externally. Send synthetic (D0) or strongly de‑identified (D1) packets to external contributors after a documented risk assessment and legal release gate; reserve any residual‑sensitive (D2) review for named, contractually restricted managed‑service work, and keep D3/D4 internal only. NIST’s SP 800‑188 explains that de‑identification reduces, rather than eliminates, risk and that protected sharing mechanisms still require assessment (NIST SP 800‑188).
Key diligence gaps to close before award: scope and timing of any security attestations, device/workspace enforcement, worker‑pay and rejection/appeals practices, contributor geography and facilities, data‑processing agreements and subprocessors, and retention and deletion evidence. The matrix below summarizes public‑fit signals and the common unknowns procurement should close with engagement‑specific evidence. It is not an endorsement or ranking.
| Provider (illustrative) | Operating model (self‑serve, recruited, expert, managed) | Public‑fit signal (to verify) | Pricing visibility | Potential MyVault fit (pilot proposal) | Main diligence gap |
|---|---|---|---|---|---|
| Clickworker | Self‑serve crowd; survey product; managed options marketed via partnerships | Product pages describe surveys and qualification filters | Survey formula public; annotation/managed work quoted (survey terms) | Sanitized, high‑volume checks on D0/D1 packets | Exact device/workspace controls; retention/deletion details; contributor locations |
| Prolific | Recruited participant panel (self‑serve studies; managed sourcing options) | Public documentation of participant targeting and recommended rewards | Transparent platform fee plus participant reward (pricing page) | Sanitized UX, comprehension, and pairwise preference studies | End‑to‑end study‑tool data path; enterprise‑security specifics by quote |
| Expert networks (various) | Screened contractors with domain credentials | Role listings and case descriptions for domain‑specific work | Quote‑only | Rubric design, senior adjudication, and staged red teaming | Credential validation, geography restrictions, secure workspace, and retention terms |
| Managed AI‑data services (various) | Dedicated teams, workflow customization, facilities | Capability marketing (annotation, RLHF, red teaming, VDI/DLP) | Quote‑only | Structured programs on D0/D1; exceptional D2 with strict controls | Attestation scope and dates; deletion proof; no‑secondary‑use commitments; subprocessors |
Any shortlist for the MyVault pilot should be framed as a set of options to test, not a predefined ranking. Select two or three vendors across models, request a comparable line‑item quote, and run a small, fully sanitized calibration to measure gold accuracy, adjudication overturns, latency, and cost per accepted judgment before expanding scope.
What Human AI Evaluation Costs — And Why Public Rates Mislead
Budgeting human review for MyVault requires three distinct price columns: public buyer list price, the amounts workers actually receive, and negotiated enterprise quotes. Conflating them is a common error; public evidence shows discrete list prices exist only for a narrow set of products and must not be generalized to secure, expert, or managed AI‑evaluation work.
1. Public buyer list prices (narrow, product‑specific evidence)
Two precise public buyer formulas from the sources above illustrate how list prices work for specific products:
- Prolific: a buyer pays the participant reward plus a platform fee (their pricing page describes a 42.8% corporate platform fee at the time of writing). That fee structure applies to Prolific studies and is not a managed‑service quote (Prolific pricing).
- Clickworker (survey product): a buyer‑set participant fee with a US$0.25 minimum plus a 40% service fee, exclusive of VAT where applicable (Clickworker survey respondents).
Neither formula is a general “AI‑evaluation” rate. They apply to the specific products described and do not establish prices for annotation, secure workspaces, managed red teaming, or enterprise programs.
Two illustrative calculations using the public formulas make the difference between participant reward and buyer invoice concrete:
- Prolific example: 1,000 judgments at three minutes each; reward set at US$12/hour. Reward per judgment = (3/60) × US$12 = US$0.60; total participant reward = US$600. Applying the published 42.8% corporate platform fee yields a buyer invoice of US$856.80 before VAT (Prolific pricing).
- Clickworker survey example: 500 completed surveys at US$0.50 each. Participant fees = US$250; 40% service fee = US$100; buyer invoice = US$350 before VAT (Clickworker survey respondents).
2. Worker pay signals (useful for fairness and capacity, not buyer invoices)
Worker compensation signals (e.g., recommended hourly rewards on participant platforms or role listings for expert cohorts) are helpful context for fairness and resourcing, but they are not buyer invoices. They vary widely by task, skill, geography, provider, and acceptance rates. Procurement should therefore request project‑specific disclosures about reviewer compensation, whether qualification and calibration are paid, and how rejected work and appeals are handled. MyVault’s pilot should treat compensation transparency as a quality and continuity control.
3. Negotiated enterprise quotes (line‑item separation required)
Most managed, expert, or secure‑evaluation work is sold by negotiated statement of work. Any RFP or SOW should present a clear, comparable price breakdown. A defensible planning model for buyer quotes separates the following elements:
| Enterprise quote component | Why it must be separated |
|---|---|
| Setup and rubric engineering | Prevents a low unit rate from masking large onboarding charges |
| Recruitment and qualification | Exposes premiums for rare experts or geography restrictions |
| Worker compensation | Enables fairness review independent of the buyer invoice |
| Platform and project management | Separates software/ops from human labor |
| Redundancy, gold, and audit sampling | Makes the cost of quality explicit |
| Expert adjudication | Prevents senior review from being silently omitted or rebilled |
| Secure workspace and geography restrictions | Identifies the privacy/security premium |
| Rework, rush, and capacity reservation | Clarifies who pays for guideline changes or minimums |
| Taxes, currency, payment fees | Allows true all‑in comparison |
Compute cost per accepted judgment as total pilot cost divided by the number of judgments that survive screening, rework, and adjudication. This denominator exposes the real cost of low quality or overturned labels. Compare providers only after equivalent quality and security controls are in place.
How Leading Labs’ Practices Can Inform the MyVault Pilot
Public documentation from large labs provides useful design patterns for evaluation — demonstrations, pairwise rankings, specialist red teaming, held‑out test sets, iterative snapshots, and layers of automated checks. These materials describe internal methods, not buyer‑facing human‑review services. They are inputs that MyVault can adapt within its governance posture and privacy constraints.
OpenAI’s instruction‑following overview, for example, discusses using demonstrations and rankings with screened labelers and cautions about the limits of average rater preference as a proxy for truth (OpenAI: instruction‑following). Anthropic’s public red‑team write‑up emphasizes that iterative oversight, re‑rating, and analysis of disagreements are necessary to reduce harms and uncover edge behaviors (Anthropic red‑team research).
These examples also highlight a boundary: lab practices are not a substitute for MyVault’s procurement. MyVault still needs contracts, sampling plans, quality gates, and privacy controls that match its specific risks and jurisdictions. NIST’s Generative AI Profile can help structure those decisions by encouraging documented evaluation roles, monitoring, and record‑keeping (NIST AI 600‑1).
- Pairwise preference helps, but is not self‑validating. Use pairwise comparisons to reduce scale‑calibration issues, but keep a separate factuality/grounding gate when truth matters (see OpenAI’s discussion of demonstrations and rankings).
- Evaluation is iterative. Convert human‑found failures into frozen regression tests and re‑run them after remediation; repeat on snapshots (consistent with patterns described in public lab materials and general risk‑management guidance).
- Experts and crowds have different jobs. Route subjective UX breadth to crowds or recruited panels on sanitized packets; handle high‑risk, technical, or legal interpretation with named experts in controlled workspaces.
- Disagreement is information. Track and analyze disagreements rather than forcing false consensus; escalate to adjudicators when needed (Anthropic’s narrative highlights re‑ratings and oversight).
To make evaluation findings durable, wire them into operations. This playbook for production model‑performance monitoring shows how evaluation can emit production signals instead of one‑off reports.
Quality Is a Measurement System, Not a Majority Vote
Human review should be designed as a layered measurement system: redundant people, guarded gold items, adjudication, and ongoing sampling produce evidence about model behavior and privacy risk. The seven‑layer control model below groups components that can be contracted, monitored, and iterated. It is a MyVault pilot proposal, not an industry standard.
Seven‑layer quality control system (pilot proposal)
- Task validity — Define the construct you measure (correctness, grounding, privacy, safety, usefulness, uncertainty). Include explicit response options for ties and cannot‑judge so ambiguity is visible rather than suppressed.
- Screening — Gate by language, domain skill, and a paid qualification so initial cohorts meet minimum competency.
- Paid calibration — Run a paid training and qualification round with feedback and documented thresholds before production; retain records of reviewer performance over time (NIST AI 600‑1 encourages retained evaluation records and assigned roles).
- Redundancy and agreement — Duplicate labels proportionally to task subjectivity and risk; measure consistency by slice and avoid treating agreement as ground truth when context is missing.
- Gold/sentinel tasks — Seed and rotate adjudicated sentinel items (easy, boundary, adversarial) to detect drift and low‑effort behavior; avoid overexposing gold items to prevent training‑to‑the‑test.
- Adjudication and audits — Escalate disagreements, privacy failures, and high‑impact items to named senior reviewers; record rationale and feed results to remediation and regression suites.
- Monitoring and iteration — Track drift, slice performance, latency, cost per accepted judgment, and incident trends; re‑run frozen holdouts after changes. NIST guidance encourages monitoring and retained records (NIST AI 600‑1).
Why inter‑reviewer consistency is not truth
Agreement measures consistency among reviewers, not factual correctness. High or low agreement can result from task ambiguity, insufficient context, divergent domain knowledge, or incentives. In the pilot, treat agreement as diagnostic: low agreement triggers rubric review, calibration, or adjudication; high agreement should still be spot‑checked against trusted references.
Why explicit uncertainty options matter
Options such as tie and cannot judge preserve measurement signals by surfacing ambiguity or insufficient context. They prevent forced choices that create an illusion of consensus and may drive incorrect retraining signals. Require brief reason codes for abstentions or ties on high‑risk tasks.
Two sampling streams: risk‑directed plus random monitoring
To balance cost and representativeness, the pilot should run two concurrent sampling streams:
- Risk/information stream: actively oversample items with low confidence, novelty, recent regressions, multilingual edge cases, or high‑impact labels. This increases the discovery rate of faults.
- Random monitoring stream: draw an unbiased sample using known inclusion probabilities so long‑term trends and population‑level error rates remain estimable (consistent with monitoring themes in NIST AI 600‑1).
Record sampling probability and reason so reported metrics can be corrected for selection bias when estimating population‑level performance.
Pilot metrics and proposed gates (for MyVault only; not industry standards):
| Metric | Definition | MyVault proposed pilot gate |
|---|---|---|
| Gold accuracy | Correct decisions on expert‑adjudicated sentinel items | ≥ 90% overall; ≥ 98% on critical privacy/permission sentinels |
| Inter‑reviewer consistency | Raw agreement plus a task‑appropriate statistic by slice | Any high‑risk slice with very low consistency triggers rubric review or pause |
| Expert‑adjudication overturn rate | Share of sampled/escalated judgments reversed by an adjudicator | ≤ 10% overall after calibration; zero tolerated for undetected critical privacy exposures |
| Privacy‑violation rate | Tasks exposing unapproved data or reproducing restricted identifiers | Zero unapproved raw‑data exposures; any exposure = immediate incident and pause |
| Re‑identification success | Share of red‑team attempts that link a packet to a real household/person | Zero successful links under approved task design (assessed per NIST SP 800‑188 concepts) |
| Disagreement/abstention | Share of items with ties or cannot‑judge | Diagnose by slice; sustained spikes prompt redesign (do not punish reviewers) |
| Latency | Time to first valid label, batch completion, adjudication, accepted export | p95 adjudication latency must fit the pilot‑defined release window |
| Cost per accepted judgment | Total pilot cost ÷ accepted judgments after rework and adjudication | Compare providers only after equivalent quality/security controls |

A Privacy‑Safe Human Review Architecture for MyVault (Pilot Proposal)
The architecture below turns “outsource evaluation, not possession” into concrete steps. It treats transformed real‑data packets as potentially personal until a documented risk assessment and legal release gate approve use.
Data classes and allowed external access
The following classification is intended to guide which MyVault items can be transformed and routed to external human reviewers during the pilot and under what constraints. These are pilot proposals, not industry standards. NIST materials stress that de‑identification reduces, rather than automatically removes, risk (NIST SP 800‑188); compliance obligations under laws like the GDPR should be reviewed with counsel (GDPR).
| Class | Definition / examples | Allowed external access (pilot proposal) |
|---|---|---|
| D0 — Synthetic | Fully fabricated family documents, invented chats, or seeded canaries created to mirror real workflows but containing no real personal data. | Allowed to self‑serve crowds, recruited participants, and expert networks after task approval. Maintain unlinkability and short retention (NIST AI 600‑1 encourages assigned oversight and retained records). |
| D1 — Strongly de‑identified | Real items transformed to remove direct identifiers and quasi‑identifiers, metadata stripped, entities tokenized, and no persistent linkage key exported; treat as potentially personal pending risk assessment. | Allowed to approved external cohorts under contract and technical controls after a documented risk assessment and legal release gate (NIST SP 800‑188). |
| D2 — Residual‑sensitive | Redacted real excerpts where contextual clues or rare combinations of attributes create residual re‑identification risk despite redaction. | Exceptional access only: named, trained, jurisdiction‑restricted expert/managed cohort in a controlled workspace with no download and short retention, after risk assessment and counsel review. Pilot default is internal only. |
| D3 — Raw personal / family data | Original documents, full chat histories, photos, medical/legal/financial records, or family‑structure mappings that directly identify people. | External access prohibited in the pilot. Internal MyVault trusted review only. |
| D4 — Secrets / control plane | Passwords, recovery codes, encryption keys, identity maps, admin logs, or any control‑plane artifacts enabling account takeover or reconstruction. | Never included in human evaluation tasks. Strict internal handling and automated monitoring only. |
Architecture: transform, test, route, adjudicate
- Local classification. Classify source items into D0–D4 inside MyVault with versioned rules and auditable decisions. Assign owners and oversight roles consistent with NIST’s governance themes (NIST AI 600‑1).
- Minimization and transformation. For D0–D2 candidates, apply field minimization, redaction of direct identifiers, tokenization of entities, cropping of images, and metadata stripping. Prefer derived outputs over raw documents.
- Automated re‑identification test. Subject transformed packets to an internal re‑identification/leakage test; reject or strengthen any packet that fails before release.
- Risk router. Route D0/D1 to self‑serve or recruited cohorts; route specialist D1 (and any exceptional D2) to a restricted expert/managed cohort in a controlled workspace; keep D3/D4 internal.
- External review on approved packets only. Workspaces receive pseudonymous task IDs with no linkage key. Enforce no download/copy/screenshot for all D1/D2 work; set short retention and require deletion attestations.
- Buyer‑controlled validation and internal adjudication. External labels enter MyVault’s validation layer (gold checks, agreement thresholds, privacy scans). Escalations and final decisions remain internal.
Minimum required controls (pilot defaults)
- Unlinkability: Token‑maps and identity linkage remain inside MyVault; no persistent mapping is exported.
- Least privilege and short‑lived access: Role‑based access with multifactor authentication, forced session expiry, and swift revocation.
- Approved countries and named cohorts: Restrict reviewer geography in contract and enforce through the workspace; disclose subprocessors.
- No download/copy/screenshot: Enforce technically for all D1/D2 work; verify during a small pilot.
- No secondary use / no training: Prohibit provider reuse of task content or rationales; reserve audit rights to verify compliance.
- Short retention and deletion attestation: Minimize retention, document deletion steps, and require deletion attestations covering caches and backups.
- Comprehensive audit logs: Record task release, reviewer pseudonym, rubric version, model version, transformations applied, adjudication, and exports.
- Incident handling and stop authority: MyVault can suspend processing immediately, require evidence preservation, and receive rapid breach notification; stop‑work clauses apply to any unapproved D2–D4 exposure.
- Worker welfare and fairness: Pay for calibration; provide clear task descriptions; allow opt‑out for distressing content; explain rejections with an appeals path. Fairness controls are both ethical and operationally stabilizing.
The MyVault Decision: A 90‑Day, Two‑Lane Pilot
The evidence and governance principles above support a practical next step: run a controlled 90‑day pilot that evaluates synthetic and strongly de‑identified artifacts rather than outsourcing raw family documents. External human reviewers operate on D0 or D1 packets only. The pilot tests whether external judgments add measurable value under auditable security, quality, and cost constraints before considering any residual‑sensitive exception.
Two lanes, different risk profiles
- Lane A — Recruited participants and broad qualified reviewers. Use a recruited participant platform for sanitized preference, comprehension, relevance, and language/UX checks. Prolific’s pricing page provides a transparent anchor for buyer calculations (participant reward plus a platform fee described at 42.8% for corporate buyers at the time of writing). Clickworker’s survey product publicly describes a US$0.25 minimum participant fee plus a 40% service fee, but this applies to the survey product and not to managed work (Prolific pricing; Clickworker survey respondents).
- Lane B — Restricted managed experts and adjudication. Reserve expert networks or managed services for rubric design, senior adjudication, and a staging‑only red team. Require line‑item quotes, no‑secondary‑use clauses, workspace controls, contributor locations, subprocessors, short retention, and deletion attestations. Run a small paid calibration to measure gold accuracy, adjudication overturns, latency, and cost per accepted judgment before any expansion.
Do not supply D3 (raw personal/family data) or D4 (secrets/control‑plane data) to any external vendor during this pilot. Any request to touch D2 (residual‑sensitive) material must follow a separate, written exception process with explicit technical and contractual mitigations and counsel review (e.g., GDPR applicability, transfer conditions, and DPIA triggers depending on facts and jurisdictions).
Phased 90‑day plan (compact)
| Phase | Days | Activities | Gate / Stop criterion |
|---|---|---|---|
| 0. Governance and scope | 1–10 | Name accountable owners; define prohibited data classes; select two use cases (e.g., sanitized RAG grounding; consent/privacy behavior). | No vendor data transfer until D0/D1 transformation and legal basis are approved. |
| 1. Evaluation asset design | 11–20 | Build 300–500 D0/D1 cases (ordinary, boundary, adversarial, multilingual). Author gold labels and rationales. Treat transformed packets as potentially personal until a release gate signs off. | Stop if any case can be linked to a real household or contains prohibited data. |
| 2. RFP & security diligence (parallel) | 11–25 | Issue a common questionnaire; obtain line‑item quotes for labor, QA, adjudication, platform, and workspace. Review certificates (scope/dates), DPAs, subprocessors, retention, locations, and no‑secondary‑use commitments. | Eliminate providers refusing location disclosure, deletion attestations, or incident/audit terms. |
| 3. Paid calibration | 26–40 | Run paid calibration with 30–50 training and 50–100 blinded qualification items per cohort; collect feedback and slice‑level gold scores; retain evaluation records (NIST AI 600‑1 governance themes). | Pause cohort if critical privacy sentinel accuracy is materially below target or overall gold accuracy is inadequate. |
| 4. Controlled production batch | 41–58 | Process 500–1,000 D0/D1 items: single review for low‑risk objective items, duplicate for subjective/high‑risk, adjudication on disagreement. Enforce workspace controls and short retention. | Immediate stop for any unapproved data exposure or re‑identification evidence. |
| 5. Expert red team & slice audit | 59–72 | Run staging‑only attacks (e.g., prompt injection, leakage); perform independent audits of random and risk‑weighted samples; analyze disagreements. | Stop if a reproducible critical privacy access issue is found or vendor prevents evidence collection. |
| 6. Remediation & re‑test | 73–82 | Fix task design/model behavior/retrieval; re‑run frozen holdout and adversarial tests without exposing previous labels. | No pass if improvements only occur on visible calibration/gold cases. |
| 7. Decision / operating model | 83–90 | Compare quality, privacy, speed, cost per accepted judgment, worker conditions, and operational fit. Prepare SOW or stop decision. | Scale only if all gates pass and accountable owners accept residual risk. |
Proposed stop criteria
- Any unapproved exposure of D2, D3, or D4 data is immediate stop‑work and incident declaration.
- Any successful re‑identification that links a packet to a real household or person is immediate stop‑work and a forensic hold.
- Evidence of unauthorized access, copying, secondary use, or undisclosed subprocessors/countries triggers contract suspension.
- Critical privacy/permission sentinel accuracy remains below target after one paid retraining cycle.
- Overall gold accuracy remains below target after calibration and feedback.
- Persistent low consistency in a high‑risk slice (after rubric fixes) pauses that task.
- Cost per accepted judgment or p95 adjudication latency materially exceeds the approved envelope.
Critical findings need an escalation path. See the related incident reporting and human stop gates guide for defining reporting, alignment gates, and stop criteria.
What to Ask Before Signing a Crowdwork or Expert‑Review Contract
This checklist converts governance principles and the pricing evidence above into an actionable RFP section for MyVault’s pilot. Items are framed as mandatory questions, evidence requests, or contractual clauses. Where public price signals exist, point to the original page so procurement can verify vendor statements.
-
Contracting and legal entity
- Which legal entity will sign the SOW, process MyVault data, and employ or contract reviewers? Provide registry evidence and a named signing officer.
- Provide a draft data‑processing agreement (DPA) and a list of subprocessors with change‑notice commitments. If GDPR or analogous regimes may apply, MyVault will consult counsel on controller‑processor roles, transfers, and DPIA triggers (GDPR).
-
Commercial breakdown and minimums
- Provide line‑item pricing for setup, rubric engineering, recruitment, platform fees, per‑judgment labor, QA sampling, adjudication, secure‑workspace, rework, rush, taxes, and currency handling.
- State any minimums, capacity reservations, and change fees. Separate setup from per‑judgment rates to avoid hidden onboarding charges.
- If citing a public price formula, show the specific bill of materials that produced any quoted total (e.g., participant reward + platform fee on Prolific; Clickworker survey fee) (Prolific pricing; Clickworker survey respondents).
-
Worker compensation, training pay, rejections, and appeals
- Report the exact amount paid to each reviewer tier (gross), whether calibration/qualification are paid, and payment timing.
- Provide written policy and SLAs for rejected work, dispute handling, and an independent appeal channel for workers.
- Declare whether buyer rates equal worker pay or whether platform/management fees are withheld; separate “worker pay” from buyer “all‑in cost.”
-
Reviewer credentials and evidence
- List identity, language, educational, license, and professional checks for the assigned cohort and attach redacted evidence samples (with hashes or attestations) that can be audited.
- Specify qualification tests, passing thresholds, and re‑qualification cadence; require paid qualification datasets.
-
Countries, facilities, and subprocessors
- Name all countries where task content, reviewer access, logs, backups, and support will occur. For any facility claim, provide current audit evidence and site‑assignment commitments.
- Disclose and contractually limit subprocessors and subcontracts; require prior notice and approval for replacements.
-
Device, workspace, and tooling controls
- State whether work occurs on provider‑managed virtual desktops/VDI, locked browsers, or personal devices; demonstrate controls (no screenshot, clipboard, print, or local download) and list enforcement mechanisms.
- Provide the stack used for task delivery and any monitoring or telemetry agents.
-
Ban on secondary use and training
- Contractually prohibit provider or subcontractor secondary use, model training, portfolio use, or onward disclosure of task content, reviewer rationales, or metadata. Reserve audit rights to verify compliance.
-
Retention, deletion, and deletion proof
- List retention windows for task data, responses, logs, caches, backups, and analytics; provide deletion procedures and a deletion attestation process.
-
Quality, metrics, and adjudication
- Specify calibration size, gold frequency, redundancy policy, expected consistency metrics by slice, and acceptable thresholds for the pilot. Require reporting of gold accuracy, adjudication overturn rate, privacy‑violation rate, disagreement/abstention rate, latency, and cost per accepted judgment.
- Describe the adjudicator role, escalation workflow, documentation required for overturned cases, and rework SLAs.
-
Incident notification, audit rights, capacity, and SLA
- Define incident types that trigger immediate notice, maximum notification windows, and rights to preserve evidence and pause processing.
- Reserve audit rights to review logs, worker qualification evidence (sampled), site controls, and task packets. State time‑to‑staff a cohort, p95 labeling and adjudication latencies, throughput guarantees, and remedies.
-
Stop‑work, termination, and remediation authority
- Grant MyVault unilateral stop‑work authority on any confirmed D2–D4 exposure, successful re‑identification, undisclosed subprocessors, or failure to report incidents. Require a remediation timeline and a termination clause for unresolved critical failures.
Worker welfare as quality control
- Require paid calibration and qualification rounds; document hours and pay bands for transparency.
- Set a fair‑pay floor and require disclosure of typical effective hourly pay for each reviewer tier; require remedial plans if actual pay drops below contracted levels.
- Provide opt‑out and support for reviewers exposed to distressing content, plus a documented appeals process for rejections.
Useful Links
- Prolific Pricing (participant reward + platform‑fee structure)
- Clickworker Survey Respondents (self‑service survey pricing)
- NIST Generative AI Profile (AI 600‑1)
- NIST SP 800‑188: De‑Identification Guidance
- GDPR Regulation (EU) 2016/679 (official text)
- OpenAI: Aligning Language Models to Follow Instructions
- Anthropic: Red Teaming Language Models to Reduce Harms
The Practical Bottom Line
The central procurement lesson is simple: a “crowd” is not a single product. Buyers face at least four operational models — self‑serve crowds, recruited participants, expert networks, and managed services — and the decisive differences are who recruits and trains reviewers, who owns rubrics and adjudication, and who controls worker access and data visibility. Public pricing examples are narrow and non‑generalizable: Prolific publicly markets a platform fee on top of participant rewards, and Clickworker’s survey offering exposes a minimum participant fee plus a service fee; both pages are useful planning inputs for sanitized, lower‑sensitivity work and are not managed‑service price lists (Prolific pricing; Clickworker survey respondents).
Outsource privacy‑preserving evaluation, not possession of the vault. For MyVault’s proposed pilot, external reviewers should normally receive synthetic cases, strongly de‑identified excerpts, or derived model outputs; raw family documents, account credentials, full chat histories, identity maps, and unrestricted product access should be prohibited by default. NIST’s de‑identification guidance underscores that risk remains and must be assessed; compliance obligations such as GDPR roles and transfer conditions depend on facts and should be reviewed with counsel (NIST SP 800‑188; GDPR).
Cost and quality are linked. Budgeting should distinguish public buyer prices, worker pay, and negotiated enterprise quotes; measure cost per accepted judgment (not per submitted label); and budget setup, calibration, redundancy, adjudication, security, and rework explicitly. Quality and privacy should be measured by defensible signals: gold‑set accuracy, adjudication overturns, privacy‑violation and re‑identification tests, latency, and cost per accepted judgment — with retained evaluation records and assigned oversight as NIST guidance encourages (NIST AI 600‑1).
Decision summary and action: run the 90‑day split‑lane pilot (recruited/participant lane + restricted managed‑expert lane), preserve MyVault’s internal authority over identity, access, consent, and safety, and treat the pilot as gated: scale only after quality, privacy, labor, cost, and operational gates pass. Cheap scale is a false economy unless it survives gold checks, adjudication, privacy controls, and worker‑welfare gates. For operational handshakes, see the related guide on human‑in‑the‑loop workflow design and the playbook for production model‑performance monitoring.
