Selina Gaertner Consulting Main site

From the talk "Compliance as Competitive Advantage", Pharma Commercial & SFE Summit, Budapest 2026

Does your AI give the same answer twice?

The 20-run test is the cheapest validation step there is. One task, twenty runs, one afternoon, no IT project. Here is how to do it, a checker to run it on this page, and a validation log template to file the result.

Twenty runs of the same task. Filled squares agree with the most common answer.

Why this matters

Language models are probabilistic. Ask the same question twice and you can get two different answers. In a consumer app that is a quirk. In pharma, where an output shapes a call plan, a CRM note or an email to a physician, it becomes a question you cannot answer later: how did the system arrive at this?

If the method does not reproduce, you cannot measure it, improve it or defend it. Practitioners at Genmab and Ipsen make the same point in a 2026 industry white paper: skills are tested for convergence before they are deployed, and measurement only works when outputs are stable enough to measure against.

Stable is not the same as correct

A system can make the same mistake twenty times. Convergence shows stability. Pair it with an accuracy check and a count of critical errors, and you have a simple validation baseline.

How to run it

  1. Pick one live task with a defined output

    A classification, an extraction, a structured summary. Free prose is hard to compare; if your task produces prose, test the fields that matter, such as next steps or recipients.

  2. Freeze the input and the version

    Same input text, same prompt or skill, same model and settings. Note the version, you will need it in the record.

  3. Run it twenty times

    Use a fresh session for each run so earlier answers cannot influence later ones. Copy every output.

  4. Check convergence

    Paste the outputs into the checker below. It groups identical answers and shows how many runs agree.

  5. Check accuracy, then file it

    Compare the answers against what is actually correct on a small sample, count critical errors, and save a dated record. That is your first validation artefact.

Run the check

The checker runs in your browser. Nothing you paste into it leaves your device, it is never transmitted to me or to anyone else. Your entries are only kept on this device if you tick the box below.

Validation record

Suggested reading of the result: 18 or more matching runs out of 20 is stable, 14 to 17 needs a closer look at the input or the output format, below 14 is not a workflow yet. Set your own thresholds for your risk level.

Get the validation log template

An Excel workbook for your runs: paste the twenty outputs, mark matches, correctness and critical errors, and it calculates convergence and accuracy for your record.

The next step: rehearse the rollout

A stable model is only one of three layers. Whether a tool survives contact with your field force, works council and compliance team depends on decisions nobody tests in advance. A multi-agent simulation lets you rehearse the rollout before you build.

If you want to talk about either, for your own use case, get in touch.

Book 30 minutes with Selina Gaertner