Skip to content
All analysis

Sector · AI · 04 Aug 2026

Label audit LA-5: the vision set that disagreed with itself

Inter-rater kappa, guideline drift, and silent relabels decide whether training truth exists—industrial AI label audit, not confusion-matrix post-mortems alone, not synthetic-data essays, not model-card rejects for soft sensors.

Label audit LA-5: the vision set that disagreed with itself

Audit: LA-5 · Vision defect label set
Sample: 1,200 images dual-labeled
Claim: “Gold labels.”
Format: Label audit. Not another accuracy slide.


Inter-rater snapshot

| Class | Agree % | Cohen-ish κ (approx) | Notes | | --- | --- | --- | --- | | Scratch | 91 | High | Guideline clear | | Stain | 74 | Med | Lighting fights | | Critical crack | 62 | Low | Worst money class | | OK | 88 | — | Default bias to OK |

Critical class is where money lives and agreement dies.

Two labelers’ boxes disagreeing on the same defect image

If humans disagree, the model is not ‘wrong’—the truth is undefined.


Findings

  1. Guideline v3 quietly replaced v2 mid-project; half the set mixed rules.
  2. “OK” used as fatigue button after hour two.
  3. Relabel by senior without version bump — training saw both truths.
  4. Synthetic overlays used crack textures that don’t match plant lighting.

Required gates before next train

| Gate | Rule | | --- | --- | | Guideline freeze ID | Must match dataset manifest | | Dual-label critical class | 100% until κ gate met | | Fatigue blocks | Max continuous label minutes | | Relabel MoC | Version + diff report | | Synthetic | Separate split, never silent merge |


Disposition

  • Freeze current “gold” for critical class.
  • Relabel critical under v4 only.
  • Confusion-matrix work waits on LA-5 close.
  • Accuracy claims blocked in customer decks until κ gate green.

Label guideline printout with version stamp v4

Unversioned guidelines are how gold turns to brass.


Adjacent fences

Confusion-matrix post-mortems own escape metrics on a deployed model. Soft-sensor model cards own process proxies. Synthetic-data essays own generation tricks. LA-5 owns whether labels are a contract. Do not retrain to chase a quarrel between two annotators.


Close

A vision model cannot outperform the agreement of its labels. Publish κ—or keep shipping “gold” that never was.

More in this sector