Ledger: RL-4 · On-prem maintenance RAG (read-only)
Sample: 40 technician questions, one week
Claim: “Answers are grounded — we show citations.”
Format: Citation failure ledger. Not a vector-DB vendor pitch.
Failure classes (counted)
| Class | Count | Example | | --- | --- | --- | | Correct cite + correct answer | 18 | — | | Correct cite + wrong synthesis | 7 | Right SOP, wrong torque | | Stale cite (retired rev) | 6 | SOP-441 Rev B cited; Rev D live | | Wrong chunk / wrong asset | 5 | Pump A procedure for Pump B | | Cite UI shown, chunk empty | 2 | Hallucinated footnote | | Should refuse, answered anyway | 2 | Safety interlock question |

A citation chip is not a verification. A revision match is.
Stale-cite autopsy (one row)
Question: “LOTO sequence for chiller-2 electrical.”
Cited: SOP-EL-19 Rev B.
Live: SOP-EL-19 Rev D (isolation point added).
Technician followed Rev B. Near-miss.
RAG “confidence” high. Ledger confidence: zero.
Ledger gates before wider rollout
- Every cite must resolve to doc ID + revision + chunk hash.
- If live rev ≠ indexed rev → hard refuse + reindex ticket.
- Safety / LOTO / SIS questions: refuse unless dual-source agree.
- Wrong-asset rate owned weekly by reliability + knowledge owner.
- Empty-chunk cites = sev-1 bug, not UI polish.

If the index lags document control, the model is a rumor engine.
Adjacent fences
Shadow diaries own operator ignore rates. Label audits own vision truth. Soft-sensor cards own process proxies. Write-path red-teams own actuation risk. RL-4 owns whether citations are contracts. Do not scale a chatbot that quotes the dead revision politely.
Close
Grounding is not a footer. Grounding is a live revision. RL-4 counted the difference—and six technicians almost paid it.
