Audit: LA-5 · Vision defect label set
Sample: 1,200 images dual-labeled
Claim: “Gold labels.â€
Format: Label audit. Not another accuracy slide.
Inter-rater snapshot
| Class | Agree % | Cohen-ish κ (approx) | Notes | | --- | --- | --- | --- | | Scratch | 91 | High | Guideline clear | | Stain | 74 | Med | Lighting fights | | Critical crack | 62 | Low | Worst money class | | OK | 88 | — | Default bias to OK |
Critical class is where money lives and agreement dies.

If humans disagree, the model is not ‘wrong’—the truth is undefined.
Findings
- Guideline v3 quietly replaced v2 mid-project; half the set mixed rules.
- “OK†used as fatigue button after hour two.
- Relabel by senior without version bump — training saw both truths.
- Synthetic overlays used crack textures that don’t match plant lighting.
Required gates before next train
| Gate | Rule | | --- | --- | | Guideline freeze ID | Must match dataset manifest | | Dual-label critical class | 100% until κ gate met | | Fatigue blocks | Max continuous label minutes | | Relabel MoC | Version + diff report | | Synthetic | Separate split, never silent merge |
Disposition
- Freeze current “gold†for critical class.
- Relabel critical under v4 only.
- Confusion-matrix work waits on LA-5 close.
- Accuracy claims blocked in customer decks until κ gate green.

Unversioned guidelines are how gold turns to brass.
Adjacent fences
Confusion-matrix post-mortems own escape metrics on a deployed model. Soft-sensor model cards own process proxies. Synthetic-data essays own generation tricks. LA-5 owns whether labels are a contract. Do not retrain to chase a quarrel between two annotators.
Close
A vision model cannot outperform the agreement of its labels. Publish κ—or keep shipping “gold†that never was.
