Every AI project has a boundary. The difference between projects is whether it was decided deliberately at the start or discovered accidentally later, usually by somebody in a compliance function reading a screen.
On the surgical decision support work in Boston, the boundary was the first thing settled. It reads the study; the surgeon makes the call. That sentence is not marketing copy. It is a design input, and almost everything else about the system follows from it.
Three tests before a model goes anywhere near the problem
We apply the same three every time, and we say no on the basis of them more often than clients expect.
- Does the decision repeat? A judgment made four times a year does not justify a model, an evaluation set or the operational cost of keeping both honest.
- Does the data already exist? If it has to be created for the project, the project is a data collection exercise with a model at the end, and it should be planned and priced as one.
- Is being wrong recoverable? This is the one that fails most often, and it is the one that decides where the boundary goes rather than whether there is one.
In clinical work the third test does not pass, and that is not a reason to walk away. It is the reason the system stops at support. The model surfaces structure and measurement; a person decides and remains accountable. Anything else would be a different regulatory object built by a different kind of company.
The boundary belongs in the architecture
A boundary that exists only in interface copy (a disclaimer under the output, a tooltip saying "for reference only") is not a boundary. It is a hope about how somebody will read a screen at the end of a long day.
A boundary in the architecture looks different. There is no code path that acts on the model’s output. There is no field the model writes that the system later treats as a decision. If a downstream feature would need one, that feature does not get built, and the reason it does not get built is written in the architecture document that a named engineer signed before the build started.
The test we use: could somebody add an "auto-approve" button next week without touching the architecture? If yes, the boundary is decoration.
What happens when the model is unsure
This is the question that separates a prototype from a system, and it is astonishing how often it has no answer. A model that has only one behaviour (answer anyway) has not been finished.
There are three honest options, and the choice between them is a product decision rather than a technical one:
- Abstain. Say nothing and show the reason. Cheapest to build, and correct far more often than teams expect.
- Escalate. Route to a person, with the context they need to decide quickly rather than a link back to the raw input.
- Fall back. Use a deterministic rule that is worse on average and predictable in the worst case.
Pick one per decision point, write it down, and build it as a real code path. A fallback that has never executed is not a fallback.
Traceability changes the data model
On imaging work, a measurement a clinician cannot resolve back to a region of a specific study is not usable, whatever its accuracy. That sounds like a UI requirement. It is not.
It means every inference has to carry a reference to its input and to the region it was derived from, and that reference has to survive storage, retrieval and display. Retrofitting that into a system that stored outputs as plain values is a rewrite of the persistence layer, which is why it goes in the architecture document rather than the backlog.
What we got wrong
Early on we treated evaluation as a phase: build it, measure it, ship it. That works exactly once. A model with no continuously maintained evaluation set degrades silently as the input distribution moves, and the first signal is a user complaint, which is the most expensive monitoring anybody has ever paid for.
Evaluation is now part of the system rather than part of the project. If there is no number somebody watches, there is no such thing as an improvement, only a change.
The version of this that applies to you
You do not need a clinical setting to need this. Any AI feature that touches money, safety, employment or a regulated decision has the same shape. Write down what the system will refuse to do, what it does when it is unsure, and who signs the release. If those three answers do not exist yet, the build has not started, regardless of what the repository says.
Written by
DevTechGuru Engineering
The engineers who built and still run the systems described on this site.
Published under the company name rather than an engineer's. That is a gap, not a house style. How we work.