The part after go-live, which is where systems are actually lost.
Instrumentation before launch rather than after the first incident, releases that can be reversed, and the same dashboards we look at.
- 01It is live
- 02Something changed
- 03Somebody noticed
- 04It went back
In productionEvery system on this site is one we still operate. The container terminal platform has no maintenance window, so each release goes out while the operation is running.
What failure after launch looks like
Most systems are lost after launch, not during the build. Nobody owns the deploy, nobody is watching, and the first sign of trouble is a phone call from a customer.
The deploy is a ritual
One person can do it, it takes an afternoon, and everybody schedules around it. So changes batch up, and each release gets riskier.
No way back
A release that cannot be rolled back is not a release. It is a commitment made on behalf of everyone downstream.
Monitoring that watches the wrong thing
CPU is green and the queue that matters has been stuck for two hours.
The bill grew and nobody can say why
Cloud spend rose 40% over a year and no single change explains it, because no change was measured.
What we hold ourselves to
Every one of these is something we do on our own platforms first.
Reversible releases
Every release goes out with a way back, including on systems with no maintenance window. The container terminal platform is the proof: releases happen while ships are moving.
Instrumentation before launch
Added before the first incident, not after it, and you get the same dashboards we do rather than a monthly report about them.
Security review
A person reviews changes that touch authentication, authorisation or data handling. Never a model alone. That line is published in full.
Scalability
Measured at the hour the system is least forgiving. Eight in the morning in Motijheel is a load profile, not a metaphor.
Cost optimisation
Reviewed against what the system is doing rather than against last month’s bill, and reported with the reason attached.
Ongoing support
A named engineer who has read the architecture, not a rota reading a runbook for the first time at three in the morning.
What we set up
None of it is exotic. All of it is the difference between a system you can change and one you are afraid of.
Cloud architecture
Sized to the load you have, with a written reason for every managed service on the bill.
Infrastructure as code
The environment is a file in your repository, not a configuration somebody clicked into a console in 2023.
CI and CD pipelines
Build, test, review gate, deploy. Repeatable and reversible, run the same way by everyone.
Containers
The same image in staging and production, so "it worked on staging" stops being a sentence anyone says.
Observability
Metrics, traces and logs against the things the business cares about, with alerts that reach a person.
Disaster recovery
A restore that has been rehearsed. An untested backup is a belief, not a backup.
What we build it on.
Grouped by what it does rather than shown as a wall of marks, and short enough that your own team can judge whether they could take it over.
- Cloud
- AWSDockerCI/CDInfrastructure as code
- Operations
- Metrics and tracingAlertingLog aggregationRelease automation
- Backend
- PythonDjangoFastAPINode.js
- Data
- PostgreSQLMySQLRedisData modelling
Questions we get asked.
The objections specific to this line, answered here rather than in a first call.
Can you take over a system somebody else built?
Usually, and the first deliverable is an honest assessment rather than a migration plan. We look at how it is deployed, what is monitored, whether a release can be reversed and where the undocumented parts are, and tell you what it costs to hold and what it costs to change.
Do you require AWS?
No. AWS is our default because it is where most of what we run already is, and the practices (infrastructure as code, reversible releases, observability against business signals) do not depend on the provider.
What does support actually mean in an engagement?
A named engineer who has read the architecture and can be asked why a call was made, an agreed response expectation, and the same dashboards we use. Not a ticket queue and not an anonymous rota.
Where to go next
You do not need to arrive with a perfect technical specification.
Start with the problem. Tell us what is slowing the business down, what you want to build, or where the current system is failing. We will tell you the shortest path forward, including when that path is not us.
Start your build