You verify an AI step the same way you verify any other system. You keep a set of cases where the correct answer is already known, and you score the output against them field by field.
A successful execution does not do this for you. It tells you the HTTP calls returned 200.
For an AI step inside a business process, the test comes down to five things:
A gold set of your own documents, 50 to 100 of them, chosen to include the awkward ones.
A typed output schema that allows null, so a missing value stays missing instead of being invented.
Separate pass rules for easy fields and hard fields, because one accuracy number hides the hard ones.
A confidence gate that routes anything below the line to a person before it is booked.
A regression run across the whole gold set after every prompt change and every model update.
Everything else is prompt tuning. Prompt tuning is not evidence.
A green execution is not a correct one
Three threads on the n8n community forum ask this almost word for word: "How do you verify AI workflows actually did the right thing?", "How do you catch workflows that run fine but do nothing?" and "How do you test a workflow edit before it touches production?".
They are circling one gap. An error trigger does not fire when a workflow writes a confident wrong number, and it does not fire when a workflow writes nothing at all. Those are the two failures that cost money, and neither one looks like a failure on the dashboard.
So the test has to compare output against a known answer, not against the absence of an exception.
Score fields, not documents
A document-level score is close to useless. If an invoice has 4 header fields and 30 line-item cells, one wrong unit price gives you 97 percent accuracy and an invoice you cannot book.
The pass rule sits on the field, and fields are graded by how much a wrong value costs. The same rule does both jobs: it scores the gold set at build time, and it decides routing at runtime.
Identity. Example: supplier VAT number, invoice number. Pass rule: Exact match, no exceptions. Runtime action on fail: Block the run, route to review.
Money. Example: total, net, VAT per rate. Pass rule: Exact match to the cent against the gold value, and net plus VAT reconciles to the total within a stated rounding tolerance. Runtime action on fail: Block the run, route to review.
Dates. Example: invoice date, due date. Pass rule: Exact match after normalising the format. Runtime action on fail: Route to review.
Line items. Example: description, quantity, unit price, line total. Pass rule: Row count must match, then cell-level match. Runtime action on fail: Route to review.
Neviox Digital CEO with an Engineer's degree in Computer Science from the University of Split, and multiple certifications, my core competencies span across front-end and back-end technologies, with an emphasis on Langchain, Next.js and AWS
Neviox Digital
Do you have a vision for a digital solution? Want to share your technical expertise or promote your brand? Let’s collaborate and build the future together!
What we shipped, what broke and what we would do differently. No news roundups.
Free text. Example: payment reference, notes. Pass rule: String similarity, not exact match. Runtime action on fail: Log, do not block.
Money fields are checked twice: once against the gold value, and once against themselves, because net plus VAT has to reconcile to the total regardless of what the model produced. The rounding tolerance matters here, because suppliers who round per line rather than per invoice produce legitimate documents that a strict sum check rejects.
The schema allows null on every field that is not structurally guaranteed, with an explicit instruction to prefer null over a guess. A null is a routed exception. An invented value is a silent loss.
Build the gold set from your documents, not from a benchmark
Published extraction benchmarks tell you how a model does on someone else's paperwork. Your suppliers are not in that sample. One n8n forum request asks for "a single invoice-classification workflow with an explicit expected label". Shared workflows almost always ship the prompt and not the expected answer.
We ask for 50 to 100 real cases, selected on purpose rather than at random:
Your top 10 suppliers or customers by volume. Why: This is most of the production traffic.
Scanned and photographed documents, not only native PDFs. Why: Photographed documents fail differently from native PDFs, and your inbox contains both.
Multi-page documents, and documents with a table that breaks across pages. Why: Line items are where extraction actually fails.
Foreign currency and a non-domestic VAT case. Why: Different field layout, different rounding.
Credit notes and corrections. Why: Negative amounts expose sign errors.
Two or three documents a new clerk would get wrong. Why: If a person needs judgement here, so does the model.
Each case gets the correct values typed out once, by someone from the team who owns the process. That is the only expensive part. It is typing, done once, by the person who already knows the right answer.
The confidence gate, and who approves what
Field-level scoring gives you a per-document decision at runtime, not just a grade at build time. Documents that pass every blocking rule go straight through. Anything that fails one, or comes in below the confidence threshold, goes to a human queue with the uncertain field highlighted. n8n has native send-and-wait-for-response operations with an approval mode, so the workflow pauses on the step that carries risk rather than on all of them.
Two numbers to agree before launch, and to keep watching after:
The share of volume that diverts to review, measured against how long a full manual entry took before. A diversion rate only matters next to the time it replaces, so agree both numbers in the same sitting.
The correction rate inside that queue. Every correction goes into the gold set, so the test gets stricter as the process runs.
The regression run after every model update
This is the part most automations skip. It is why a workflow that passed at launch quietly stops passing. Model versions change. Prompts get edited. Neither event announces itself in your output.
n8n ships the machinery for this. Test cases live in a Data Table. An Evaluation Trigger runs the workflow over every row. The Evaluation node writes results back with Set Outputs and records scores with Set Metrics, using built-in metrics including correctness, string similarity and categorisation. Results roll up per run in the Evaluations tab, so two runs can be compared.
Pick the instrument to match the field. Correctness is an AI judge on a 1 to 5 scale, which is fine for free text and useless for a total. Identity and money fields get a plain equality check in a Code node, not a judged score.
What matters more than the tooling is the trigger policy. We re-run the full set before any prompt change ships, when the provider releases a new model version, and monthly regardless. A drop on one field class is visible immediately, and the gold set is what lets us move the step to a different model or provider without taking the result on faith.
What this costs to run
The build cost is visible. These four are not:
Token cost per document, which scales with page count and rises when you send page images instead of extracted text. Price both paths on ten of your own documents before you choose.
Re-running the gold set. 100 documents per regression run, several runs a month.
The review queue. Real minutes from a real person, at whatever rate the diversion lands on.
One cost ceiling per run. Check whether your n8n version caps spend per execution. If it does not, a loop on a failing step has no upper bound until you add one. We set it explicitly.
We run this on self-hosted n8n in the EU, or on the client's own servers, so the gold set and the execution data stay inside the business. The document still leaves for whichever model reads it, unless the model runs locally too. That choice is made per process, in writing. How we scope and build these workflows is on our page.
Who signs off
The person who owns the process, not the person who built the workflow. They see the field-level results on their own documents, they see which cases still route to review, and they decide whether that is good enough to go live.
Before that conversation, run the gold set past the person doing the job today. Record their field-level error rate on the same documents. That is the number the automation is actually competing with, and it is almost never zero. Comparing a model against perfection is how a working automation gets rejected and a manual process with the same error rate gets kept.
If nobody is willing to put their name on the result, the automation is not ready, and no accuracy percentage will change it.
If you have a process you are thinking about automating and want to know what the test set would look like for it, send us the three documents that give your team the most trouble and we will tell you what we would measure.
Sources
How do you verify AI workflows actually did the right thing?: https://community.n8n.io/t/how-do-you-verify-ai-workflows-actually-did-the-right-thing/304241
How do you catch workflows that run fine but do nothing?: https://community.n8n.io/t/how-do-you-catch-workflows-that-run-fine-but-do-nothing/308708
How do you test a workflow edit before it touches production?: https://community.n8n.io/t/how-do-you-test-a-workflow-edit-before-it-touches-production/313427
a single invoice-classification workflow with an explicit expected label: https://community.n8n.io/t/looking-for-1-invoice-classification-workflow-with-an-explicit-expected-label/284760