Section 035 · Chapter 5, Judges, Humans, and Disagreement

Data Labeling Dangers and Labeler Demographics

The people and systems that create labels become part of the product's definition of quality.

release gatedata labeling dangers labeler demographics

What to do

  1. Use overlap and agreement metrics for ambiguous cases.
  2. Use experts for high-risk domains, calibration sets, severe failures, and labels that define release gates.
  3. Use LLM labelers for scale only after calibrating them against humans who actually understand the domain.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for release gate, data labeling dangers labeler demographics needed to reproduce work on Data Labeling Dangers and Labeler Demographics.
  • Report results for release gate, data labeling dangers labeler demographics by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

The deeper move is to treat labels as evidence with provenance, not as truth. Every important label should have a source: who or what produced it, under which guideline, with which expertise, in which context, at what time, with what disagreement, and with what adjudication path.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Data Labeling Dangers and Labeler Demographics." Testing AI Knowledge Edition, section 35.

https://jarbon.ai/testing-ai/knowledge/ch035-data-labeling-dangers-labeler-demographics.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400