gabrielwang.ai ← back to the workshop
Edition 017 · interactive demo

The findings are identical. One of these is usable.

Automation makes producing a report almost free, which quietly moves the whole problem downstream: the document now has to be acted on by a named colleague, this week, on a phone. Most generated reports are still shaped for a management review — score tiles, a chart, an overall assessment — with the only actionable content in an appendix. Below is one verification run laid out both ways. Open the three things you must sort out, in each.

Honest-AI note. The findings come from a real class of work — cross-checking a week of operational figures against a weekly summary — but the operation, the week and every number here are fictional. No model runs on this page. The interesting checks are the boring ones, and they run live: the coverage arithmetic has to foot, findings are grouped by cause so the count reads as things to fix rather than accusations, derived figures are labelled as ours, and a banned-phrase check runs over every string either document emits.

Synthetic findings Distance measured, not asserted Runs in your browser No live model call

Struck out — and why

Plain-language key (cell, root cause, derived figure, recorded-not-judged, coverage)
Cell
One number in one place. A single mistake upstream can make several cells disagree, which is why cells are a bad unit for counting mistakes.
Root cause
The one thing that has to be sorted out. Findings are grouped by it so the reader sees three tasks, not ten complaints.
Derived figure
A number we calculated rather than read. It belongs in the body with its arithmetic shown — never in the column of figures the reader published.
Recorded, no view formed
Something copied down where there is no second source to compare it against. Saying it plainly avoids implying a fault that was never found.
Coverage
How much was actually checked. Quote it only if you counted it, and show the parts adding to the total.