Diary: SD-11 · Quality assist model (read-only recommendations)
Line: Visual + process hint (anonymized)
Claim: “Shadowing for two weeks before write-path.”
Format: Day-by-day operator diary. Not an MLOps checklist.
Day 1
Model flagged 14 “review” units. Operators opened 3.
Reason logged: “busy.”
Shadow truth: ignore rate is already a KPI — nobody named it.
Day 3
Agreement when opened: 5 / 8.
Two accepts were wrong in lab (false confidence).
No owner for false-accept review. Model kept the same threshold overnight.

Unread recommendations are not data. They are wallpaper.
Day 5 — the quiet failure
Model suggested hold on a lot. Supervisor overrode verbally.
MES shows ship. Diary shows no reject code for the override.
Shadow mode cannot learn from invisible overrides.
Day 7 — scoreboard (honest)
| Metric | Value | | --- | --- | | Recommendations shown | 96 | | Opened | 31 | | Accepted | 12 | | Rejected with reason | 9 | | Ignored / no action | 65 | | Overrides logged in MES | 2 |
Shadow theater score: high activity, low learning.

If ignore has no column, you are not in shadow mode — you are in denial.
Gate before any write-path talk
- Ignore rate owned by operations (target + review).
- Every override → coded reason in MES.
- False accept / false reject weekly adjudication (human).
- Threshold changes via MoC, not night courage.
- Confusion-matrix / label work may proceed — write-path still blocked.
Adjacent fences
Confusion-matrix post-mortems own escape math on a deployed scorer. Label audits own training truth. Soft-sensor model cards own process proxies. Red-team write-path reports own dangerous actuation. SD-11 owns whether shadow is a contract with operators. Do not promote a model that spent a week talking to itself.
Close
Shadow mode is not quieter production. It is louder accounting. SD-11 proved the accounting was optional—so the learning was too.
