AI that reaches production, and stops where it should.

Most AI work fails between the demonstration and the deployment. We build the part after the demo: the boundary, the evaluation, the fallback, and the person who signs.

  1. 01Where AI fits
  2. 02Where it stops
  3. 03What ships
  4. 04Who is accountable

In productionComputer vision over DICOM studies with an orthopaedic practice in Boston, supporting preoperative planning and intraoperative guidance.

The opportunity, and the three ways it goes wrong

A model belongs where the decision repeats, the data already exists, and being wrong is recoverable. Where those three are not all true we will say so, and the answer is usually ordinary software.

  • The pilot that never shipped

    It worked in a notebook on last quarter’s data. Nobody could say what it should do when it is unsure, so it stayed a notebook.

  • Nobody drew the line

    The system quietly makes a decision somebody is legally accountable for, and the first time anyone notices is during a review.

  • No way to tell if it got worse

    A model with no evaluation set degrades silently. The first signal is a customer complaint, which is the most expensive monitoring available.

  • AI where software would do

    A deterministic rule was available, cheaper, auditable and correct, and a model was used because a model was the ask.

What production actually requires

The gap between a working prototype and a system a business depends on is almost entirely made of these.

  1. Data pipelines

    Where the data comes from, how often, what happens when it stops, and who is told. Most model failures in production are data failures wearing a model’s clothes.

  2. Model integration

    The model is a component in a system, not the system. What the surrounding software does when the model is slow, unavailable or unsure is the design.

  3. Evaluation

    A held-out set that reflects the real distribution, and a number you can watch move. Without it there is no such thing as an improvement, only a change.

  4. AI architecture

    Written down and signed before the build, including what the system will refuse to do. That refusal is the most important line in the document.

  5. Production deployment

    Versioned models, reversible releases and a fallback that is a real code path rather than an error page.

  6. Security and governance

    What leaves your network, what is retained, what a prompt may reach. Answered before a security reviewer asks, and published on our own trust page.

What we build

Six shapes, and the interesting question about each one is where the human stays.

  • Generative AI features

    Drafting, summarising and extraction inside a product, with the output treated as a draft rather than as an answer.

  • AI agents

    Multi-step tools that act. Scoped by what they are allowed to touch, not by what they are asked to do.

  • Retrieval augmented generation

    Answers grounded in your own documents, with the citation as a hard requirement rather than a feature.

  • Computer vision

    Reading images where a human is already reading them and the volume is the problem. See the AvailOrtho engagement.

  • Machine learning

    Prediction and classification where the decision repeats often enough to be worth modelling and cheap enough to be wrong about.

  • Intelligent automation

    Removing a repeated manual step, with a defined behaviour for the case the model has not seen.

OutcomeAI in the parts of the workflow that can carry it, and a written line around the parts that cannot.
CapabilitiesAI agents · LLM applications · Retrieval augmented generation · Machine learning · Computer vision · Intelligent automation

What we build it on.

Grouped by what it does rather than shown as a wall of marks, and short enough that your own team can judge whether they could take it over.

AI
LLMsRetrieval augmented generationComputer visionMachine learning
Backend
PythonDjangoFastAPINode.js
Data
PostgreSQLMySQLRedisData modelling
Cloud
AWSDockerCI/CDInfrastructure as code
Operations
Metrics and tracingAlertingLog aggregationRelease automation

Questions we get asked.

The objections specific to this line, answered here rather than in a first call.

How do you decide whether a problem should use AI at all?

Three tests. The decision has to repeat often enough to be worth modelling, the data has to already exist rather than needing to be created, and being wrong has to be recoverable. If any of the three fails we say so, and the honest answer is usually deterministic software.

Will the model make decisions our regulator holds us responsible for?

Not unless you ask for it in writing and it is lawful where you operate. Our default is that the system presents and a person decides. On the surgical decision support work in Boston that boundary is in the architecture, not in the interface copy.

What happens when the model is unsure?

It is a designed code path, decided before the build: abstain, escalate to a person, or fall back to a deterministic rule. A system whose only behaviour under uncertainty is to answer anyway has not been finished.

Does our data get used to train anything?

No, and where a third-party provider is in the path we tell you which one, what reaches it and what it retains, before the build rather than during a security review.

You do not need to arrive with a perfect technical specification.

Start with the problem. Tell us what is slowing the business down, what you want to build, or where the current system is failing. We will tell you the shortest path forward, including when that path is not us.

Start your build