The first question we ask about a new system is not what it should do. It is what it is not allowed to do. On the terminal and port platform the answer came back in four words: it cannot stop.
A container terminal runs continuously. Ships arrive on a schedule set by tide and berth availability rather than by anybody’s deployment window, gates move, equipment works. There is no two-hour Sunday morning where the operation pauses and somebody runs a migration. That is not a difficult version of a normal problem. It is a different problem, and it changes what "done" means.
Uptime is the wrong target
The instinct when somebody says "it cannot stop" is to chase availability: more replicas, faster failover, a higher number of nines. That is worth doing and it is not the thing that saves you.
A system with excellent uptime and a bad release is still down, and it is down in the worst possible way: running, accepting writes, and wrong. Recovering from that costs far more than recovering from an outage, because an outage stops the damage and a silent corruption does not.
So the property we actually optimise for is reversibility: how long it takes, and how confident we are, that we can put the system back the way it was five minutes ago with the data intact.
What reversibility costs
It is not free, and most of the cost lands on the schema rather than on the deployment pipeline.
- A migration has to be able to run while the previous version of the application is still serving traffic. That rules out anything that renames or drops in one step.
- Every schema change becomes at least two deployments: add the new shape, backfill, switch the reads, and only then remove the old one, with a stop point between each.
- The application has to tolerate both shapes at once for as long as the rollback window lasts, which means the code carries a branch nobody wants and everybody has to remember to delete.
- The rollback has to be rehearsed. An untested rollback is a belief, and the moment you need one is the worst moment to find out which.
That is roughly a 20–30% overhead on schema-touching work. We do not have a measured figure to publish for it, and we are not going to invent one. What we can say is that it is a real, visible cost that shows up in every estimate, and that it is the cheapest insurance available on a system where being wrong for an hour is more expensive than being down for ten minutes.
The mistake we made
The first version of our deployment process treated the rollback as a database restore. It was documented, it was correct, and it was useless.
A restore takes the data back to the snapshot. In a continuously-writing operation, the snapshot is minutes old and those minutes contain real container moves that physically happened. Restoring throws them away, and the yard does not restore with the database. You end up with a system that is internally consistent and disagrees with the actual world, which is worse than the state you were trying to escape.
The fix was to stop treating data and code as one artefact. Code rolls back in seconds. Data rolls forward, always, and the schema is designed so that the previous version of the code can still read what the new version wrote. That is the constraint that makes the two-step migration non-negotiable, and we learned it the expensive way rather than from a book.
What this looks like in practice
- Deployments are boring, frequent and small. A large release is a large rollback surface, and reversibility gets worse as batch size grows.
- Feature flags carry behaviour changes so that the deploy and the change are separate events with separate blast radii.
- Every release has a named approver who can be asked why they approved it. That is a person, not a pipeline stage.
- The alert that matters is not CPU. It is the operational signal (moves recorded, gate events processed), because that is the number that tells you the system is wrong rather than slow.
Why this generalises
Very few businesses have a terminal’s constraint in its pure form. Most have a weaker version of it: a plant that runs two shifts, a retail chain whose quiet hour is three in the morning in one time zone and lunchtime in another, a municipal system whose peak is eight o’clock and whose failure is publicly visible.
The useful move is to ask the terminal question of every system: if this release is wrong, what is the fastest honest way back, and has anybody tried it? A team that can answer both halves has a different relationship with shipping than one that cannot, and it shows up long before an incident does.
Written by
DevTechGuru Engineering
The engineers who built and still run the systems described on this site.
Published under the company name rather than an engineer's. That is a gap, not a house style. How we work.