If people review your AI's recommendations, a pilot shows whether they form their own judgment first, or mostly agree with the machine. You get an evidence report you can put in the audit file.
Rules and policies increasingly expect effective human oversight of AI, and they name automation bias as the risk to guard against. But a review step can look fine on paper while reviewers approve whatever the AI suggests. Nothing in a normal audit trail tells the two apart.
The sealed step tells them apart. Before the AI's recommendation is shown, the reviewer records their own call. Then it's revealed, and they submit a final decision with a reason for any change.
Later, QA records the correct decision from appeals, audits, or chart review. Reviewers never grade their own work.
The full report adds attestation and documentation rates, accuracy of the sealed call, the AI and the final decision, a per-reviewer table framed for coaching, the limits of the evidence, and a recommendation. Raw evidence exports as a CSV.
The first pilots are free in exchange for feedback. Send a short note with the workflow you'd test, roughly how many reviews a week, and who would enter outcomes.
drplummer@gmail.com
Pilots are run by someone independent of the workflow being assessed.