Production planners have always estimated changeovers. The new failure mode is quieter: a scheduling model or assistant fills setup blocks from sparse MES timestamps, averages unlike product pairs, or copies a “standard 45 minutes” that nobody timed with a stopwatch. The Gantt chart turns green. The shift starts late on the second SKU and never recovers.
This is not about prompt injection into maintenance tickets, and it is not a generic lecture on digital twins. It is specifically changeover duration as a data problem when software is allowed to invent the minutes that sequencing depends on.
What bad duration data looks like
MES often stores order start/stop. It rarely stores a clean “setup begin / setup ready” pair. People use downtime codes inconsistently. A wash that took two hours because of a blocked drain becomes the “normal” time for that product family. A lightning-fast changeover done by the plant’s best crew becomes the optimistic default the optimizer prefers.
Pairwise matrices (from SKU A to SKU B) are worse when half the cells are empty and the model imputes them. Imputation is not measurement.
What I would trust instead

- Timed studies for the high-runner transitions, with a written start/finish definition (last good part to first good part, including QC release if that gates the line).
- Separate clean, mechanical, and recipe/download steps so a slow wash does not inflate a tooling swap.
- Ranges or percentiles in the scheduler—not a single fake precision minute—when the sample size is small.
- A human gate before the optimizer is allowed to use any auto-estimated cell.
Illustrative scenario (not a named plant claim)
Suppose a scheduler assumes 30 minutes for every film change on a pouch line because three historical orders happened to be near that. On a Monday with a full sanitation requirement between allergen SKUs, the real gap is closer to ninety minutes. The plan still shows six changeovers. By noon the sequence is fiction. Nobody lied in the UI. The duration table did.
If your site uses learning-based estimates, keep the training labels auditable and refuse to deploy cells with fewer than a minimum number of clean observations. A thinner schedule with honest gaps beats a dense schedule that shifts cannot run.
I am more convinced by a short matrix of measured transitions than by a confident model that never shows its sample sizes. Where the data are thin, say so on the planning board.
