The part after go-live, which is where systems are actually lost.

Instrumentation before launch rather than after the first incident, releases that can be reversed, and the same dashboards we look at.

  1. 01It is live
  2. 02Something changed
  3. 03Somebody noticed
  4. 04It went back

In productionEvery system on this site is one we still operate. The container terminal platform has no maintenance window, so each release goes out while the operation is running.

What failure after launch looks like

Most systems are lost after launch, not during the build. Nobody owns the deploy, nobody is watching, and the first sign of trouble is a phone call from a customer.

  • The deploy is a ritual

    One person can do it, it takes an afternoon, and everybody schedules around it. So changes batch up, and each release gets riskier.

  • No way back

    A release that cannot be rolled back is not a release. It is a commitment made on behalf of everyone downstream.

  • Monitoring that watches the wrong thing

    CPU is green and the queue that matters has been stuck for two hours.

  • The bill grew and nobody can say why

    Cloud spend rose 40% over a year and no single change explains it, because no change was measured.

What we hold ourselves to

Every one of these is something we do on our own platforms first.

  1. Reversible releases

    Every release goes out with a way back, including on systems with no maintenance window. The container terminal platform is the proof: releases happen while ships are moving.

  2. Instrumentation before launch

    Added before the first incident, not after it, and you get the same dashboards we do rather than a monthly report about them.

  3. Security review

    A person reviews changes that touch authentication, authorisation or data handling. Never a model alone. That line is published in full.

  4. Scalability

    Measured at the hour the system is least forgiving. Eight in the morning in Motijheel is a load profile, not a metaphor.

  5. Cost optimisation

    Reviewed against what the system is doing rather than against last month’s bill, and reported with the reason attached.

  6. Ongoing support

    A named engineer who has read the architecture, not a rota reading a runbook for the first time at three in the morning.

What we set up

None of it is exotic. All of it is the difference between a system you can change and one you are afraid of.

  • Cloud architecture

    Sized to the load you have, with a written reason for every managed service on the bill.

  • Infrastructure as code

    The environment is a file in your repository, not a configuration somebody clicked into a console in 2023.

  • CI and CD pipelines

    Build, test, review gate, deploy. Repeatable and reversible, run the same way by everyone.

  • Containers

    The same image in staging and production, so "it worked on staging" stops being a sentence anyone says.

  • Observability

    Metrics, traces and logs against the things the business cares about, with alerts that reach a person.

  • Disaster recovery

    A restore that has been rehearsed. An untested backup is a belief, not a backup.

OutcomeA system you can change on a Tuesday without holding your breath.
CapabilitiesCloud architecture and migration · CI and CD pipelines · Observability and alerting · Performance and cost tuning · Security review · Ongoing engineering support

What we build it on.

Grouped by what it does rather than shown as a wall of marks, and short enough that your own team can judge whether they could take it over.

Cloud
AWSDockerCI/CDInfrastructure as code
Operations
Metrics and tracingAlertingLog aggregationRelease automation
Backend
PythonDjangoFastAPINode.js
Data
PostgreSQLMySQLRedisData modelling

Questions we get asked.

The objections specific to this line, answered here rather than in a first call.

Can you take over a system somebody else built?

Usually, and the first deliverable is an honest assessment rather than a migration plan. We look at how it is deployed, what is monitored, whether a release can be reversed and where the undocumented parts are, and tell you what it costs to hold and what it costs to change.

Do you require AWS?

No. AWS is our default because it is where most of what we run already is, and the practices (infrastructure as code, reversible releases, observability against business signals) do not depend on the provider.

What does support actually mean in an engagement?

A named engineer who has read the architecture and can be asked why a call was made, an agreed response expectation, and the same dashboards we use. Not a ticket queue and not an anonymous rota.

You do not need to arrive with a perfect technical specification.

Start with the problem. Tell us what is slowing the business down, what you want to build, or where the current system is failing. We will tell you the shortest path forward, including when that path is not us.

Start your build