Skip to content

An evaluation set makes “any model” true

Since the rebuild I say Inora is tied to no single language model. The sentence is only true if we can show it. Without the organisation’s own questions and answers, swapping a model is a leap of faith; with them it is a test run. Why I count evaluation as a precondition, and where I part ways with the standard advice.

Written by

Essay3 min read

Since the rebuild I keep saying that Inora must not depend on one language model. It is an easy sentence. This week I asked what I would need in order to prove it, and the answer is less comfortable than the sentence.

A swap without a test is a leap

Eleven weeks after the first organisations started working with Inora, the choice of model has become a variable for us instead of a given. A new model comes out, someone says it is better, and then what? Better at what, and on whose questions? A public leaderboard says how a model does on other people’s tasks. It says nothing about a night-shift question on a wound protocol that exists in two versions.

So I now count the organisation’s own set of questions, with the answers its own quality staff consider good, as a precondition of “any model”. Without it, a swap is a leap of faith. With it, a swap is a run: the same questions, the new model, and a comparison. Anyone who promises freedom of choice without a way to check it is promising something they cannot see.

None of this is new, and I did not think of it. Hamel Husain argued that products built on language models tend to fail for one shared reason: nobody wrote product-specific tests, and the cure is to turn every failure you see into a test that stays. Sayash Kapoor and Arvind Narayanan showed how often agents are judged on accuracy alone, with cost left out, and asked for a comparison with a cheap baseline at equal cost. I take both over as they stand.

What counts as a good answer

For us the errors are not symmetrical. An answer that sounds sure while the sources say nothing is worse than a plain “this is not in our sources”. The second counts as a correct result. The reference answer is the organisation’s own protocol, and if a model is ever used to judge another model, it has to be aligned with the person who owns that protocol. Not with me, and not with a vendor.

Where I disagree

Sierra, which sells platforms for customer-facing agents, describes in The Agent Development Life Cycle versioned releases and regression suites. I like both. What I distrust is knowledge that ships frozen inside a release, because then a superseded protocol travels along with the agent that cites it. Knowledge needs its own clock, and every answer should record which version it rested on.

Testing on care workers, the way a web shop tests two buttons, is out of the question. A new model first runs alongside the old one on the same questions while nobody sees the result, and only then reaches a small group.

Traces of real questions are part of care records. An evaluation that needs them to leave the organisation to be run elsewhere defeats its own purpose. It runs where the data lives.

And fine-tuning as the way to make a model fit: it ties you to the weights of one model. What you keep when the model changes is the set of questions and the sources with their owners.

What I am not saying

I am not reporting results. We have been live for eleven weeks, and what I describe here is how I want to work, and what I want in place before the first swap. The sentence “any model” is a promise. The evaluation set is how you keep it.

Work in care and use Inora? Sign-in and help go through your own organisation: your team’s project lead and ambassadors can help you.