Most vision and soft-sensor reviews still pick a threshold from a ROC or a confusion matrix that pretends every error weighs one. Plants do not. A false reject on a high-value module can mean tear-down, retest, and lost takt. An escape can mean a field return, a line stop at the customer, or a regulatory notice. Those numbers are not interchangeable, and an F1 that looks “balanced†can still be the wrong operating point.
This is not score-calibration-vs-base-rate, not active-learning queue bias, and not recipe MoC freezes. It is threshold economics: putting money next to the cut.
What changes when you price the errors
Define, per class or station: cost of false reject (scrap or rework + capacity), cost of escape (containment + warranty + reputation proxy you are willing to write down), and volume. The optimal threshold is the one that minimizes expected cost—not the one that maximizes F1 on last month’s holdout.
When escape cost is 50× reject cost, you will live farther into the high-recall region than a textbook ROC suggests. When reject cost is dominated by a scarce rework bay, you will tighten differently. Write the assumptions; do not hide them in a slide titled “AI quality.â€

What to show in the review pack
- Cost-weighted expected loss vs. threshold curve for the current mix.
- Sensitivity: what happens if escape cost is wrong by 2×.
- Volume at the proposed cut: how many units hit rework per shift.
If finance will not sponsor numbers, quality should still refuse a threshold that only cites F1. A placeholder cost table with named owners beats a silent zero.
Gate before promotion
No champion threshold without a dated cost table and a named owner for each cost line. Recheck after mix or ASP changes—the same probability cut is a different business decision when the product mix flips.
The model outputs scores. The plant pays cash. Put cash on the axis before you freeze the cut.
