Active learning is sold as thrift: label less, learn more. On a plant floor it often becomes thrift of a different kind—reviewers clear the high-confidence queue because it is fast, and the ambiguous frames sit in a backlog that never closes. The next train inherits another week of easy agrees. Hard lighting, rare alloys, and overlapping defects stay unlabeled. The model gets better at what it already knew.
This is not look-ahead feature joins, not scrap-code taxonomy merges, and not operator-override telemetry. It is selection bias in the human labeling loop.
What the queue is selecting
Confidence-sorted queues optimize for agreement speed. Disagreement and low-confidence items are exactly the mass that moves decision boundaries. If the KPI for the labeling team is “items closed per hour,” the system will invent a curriculum of easy scrap and call it active learning.
A healthy loop forces a quota of low-confidence and disagreement samples into every batch—before the green pile is touched.

Gates before the next retrain
Report, per train: fraction of labels from low-confidence buckets, class coverage of rare defects, and age of the oldest unlabeled hard case. Refuse a champion promotion when the hard-case backlog grew while easy-case throughput looked excellent.
Keep a frozen challenge set of known-hard frames that must not improve only because they were never scored in review.
What quality meetings should ask
Before celebrating a jump in F1 after an “active learning sprint,” ask how many borderline frames were labeled versus auto-accepted. If the answer is mostly green, you did not buy information—you bought confirmation.
The short rule: active learning without a hard-case quota is confirmation bias with a smarter sort order.
