Skip to main content
Technology Blog

Zero-Downtime Deployment: What It Requires in Practice

a laptop in a dark room showing a code editor

Zero downtime is one of those phrases that gets agreed in a meeting without anyone defining what it means, and then gets discovered to mean something quite different by each party involved. The sales conversation hears no user-visible interruption. The engineer hears a rolling replacement with health checks and a database strategy. The support team hears that nobody needs to be told, and that no call should arrive.

All three are achievable. None of them is achievable on every system. The useful work happens before the promise is written down, when you work out which of them the system can actually deliver and what it will cost to get there.

What zero downtime actually means

Define it in measurable terms, because the vague version cannot be tested and therefore cannot be delivered. Reasonable definitions include: no request to the service returns an error during the deployment window; the number of requests failing stays within the normal background rate; or specific critical user journeys are never interrupted while other parts may be briefly unavailable.

a spiral notebook and pen laid out on a wooden desk

That third definition is the most honest for most business systems and the least common. A reporting page can be unavailable for ninety seconds during a release and nobody will care. A payment submission cannot, and the difference is worth making explicit in the agreement rather than discovering it during the first incident.

Write down what happens to in-flight requests as well. Replacing an instance terminates its connections, and a user mid-submission may see an error even though the service is technically healthy a second later. Whether that counts as downtime depends on the user journey, which is exactly the kind of thing that is invisible until you ask.

Also define the ceiling. Some systems cannot achieve true continuity without cost that makes the work pointless: an on-premise installation at a customer’s site, a device with a fixed maintenance window, a regulatory environment where changes are approved for a specific date. For those, the honest deliverable is a shortened maintenance window with a rehearsed procedure, and saying so early is far more useful than attempting the impossible and missing a target.

The database decides how hard this is

Application code is the easy half. You can run two versions of a stateless service side by side without difficulty, which is the entire basis of most zero-downtime approaches. The database has no such option, because there is only one copy of the data and both versions of the code will be talking to it during any overlap period.

a spiral notebook and pen laid out on a wooden desk

So the constraint is: while old and new code run simultaneously, the schema and the data must satisfy both. Most downtime incidents attributed to deployments are actually schema incompatibilities, and they happen because the change was designed from the code’s point of view and the data’s point of view was considered afterwards.

Several categories of change are more dangerous than they look:

  • Renaming a column or table, because the old code still selects the old name.
  • Removing a column, because the old code may still select it and will start erroring.
  • Changing a column type, because reads and writes may both need different handling during the transition.
  • Adding a constraint that existing data does not satisfy, which fails at an unpredictable moment rather than at deploy time.
  • Adding an index without the concurrent option on a large table, which locks writes for the duration.
  • Any change where the default value differs between what the old code expects and what the new code writes.

Sequential migrations make this worse. If migration tooling runs one statement at a time and later statements assume earlier ones completed, a failure halfway leaves the schema in a state that neither version of the application fully understands. On a system where a deploy can be rolled back, that half-applied state is the thing that prevents the rollback.

The practical conclusion is that the database change needs to be designed as its own deployable step, tested against a realistic copy of production data, and rehearsed separately from the application release. We have found it useful to treat schema change as a separate item in the release plan with its own approval, even when it ships minutes earlier.

Schema changes: expand and contract

The pattern that resolves most of this is called expand and contract, and it is worth understanding even if you implement it by hand rather than through a migration framework.

a spiral notebook and pen laid out on a wooden desk
  1. Expand. Add the new structure alongside the old one. Nothing reads it yet. This change is backwards compatible: old code is unaffected, and new code is not using it either.
  2. Migrate. Backfill the data in batches, with the application writing to both structures so no row is missed. This is the long step and it is safe to stop partway through.
  3. Switch. Change the application to read from the new structure and stop writing to the old one. From here on, the previous version of the code is no longer compatible with the schema.
  4. Contract. Remove the old structure, in a later deploy, once you are confident nothing needs to go back.

Three or four deploys instead of one, and each one is individually reversible except the switch. That single irreversible step is where you need the most care, and it is also where knowing your rollback options is worth the most.

Backfilling deserves a note of its own because it is where these projects stall. A large table cannot be updated in one statement without locking or exhausting the transaction log, so the update needs to run in batches, with a pause between them, and it needs to be restartable after a failure. Writing that logic properly takes longer than the deploy itself. Doing it naively on a production-sized table is the single most reliable way to turn a routine release into an incident, and it is a common inheritance on projects we look at during legacy modernisation work.

Choosing a rollout strategy

Three approaches cover most situations, and the differences are about blast radius rather than difficulty.

  • Blue and green. Two identical environments, deploy to the idle one, verify, then switch traffic. Rollback is a route change and takes seconds. The costs are double the infrastructure while both run, a data synchronisation problem if the two environments share a database, and cache or session complications. It suits systems where a fast, rehearsed rollback matters more than the extra spend.
  • Canary. Deploy to a small slice of traffic, watch, then widen. The smallest blast radius of the three, and it catches problems that only appear under real load or with real data. It requires the ability to route a percentage of traffic and to compare behaviour between versions, and it needs patience, because a small slice may not surface the problem you are watching for during the window you can afford to wait.
  • Rolling. Replace instances in batches. Cheapest, no extra environment, but partial failure is possible: some instances on the new version and some on the old, which is exactly the overlap the database work has to tolerate. Rollback is a batch operation rather than an instant switch.

For a single-instance application with a shared database, none of these gives you anything much, because there is only one thing to replace and it has to stop at some point. In that situation the honest answer is that you are looking at a very short outage, and the work worth doing is making it short and making it routine.

Choosing one and writing it down matters more than the choice itself. Mixed mental models across a team produce deployments that are improvised, and improvised deployments are where the risk actually lives. The mechanics of running any of these approaches sit within CI/CD deployment work, and SmartEdge IT Solutions will usually write the chosen strategy into the delivery documentation rather than leave it in someone’s head.

Health checks that tell the truth

A traffic switch driven by a health check is only as good as that check, and shallow checks are the most common failure in this whole area. A route that returns the right status code when the process is up says almost nothing: the process can be up while it cannot reach the database, cannot resolve a configuration value, or is serving requests that all fail downstream.

a spiral notebook and pen laid out on a wooden desk

Useful checks separate two concerns. Liveness answers whether the process should be restarted, and it should be cheap and dependency-free, because a liveness check that fails on a database blip will restart healthy instances during someone else’s outage. Readiness answers whether the instance should receive traffic, and it may include dependencies, because that is precisely the question.

The other requirement is time. A check that returns immediately on the first failure will pull every instance out of the load balancer during a brief network hiccup and create the outage it was supposed to prevent. Readiness needs a threshold, and so does automatic recovery, and both thresholds should be tuned against what your service actually does rather than left at defaults.

A useful discipline is verifying the health check catches things. Take the database away in staging and confirm the check fails. Break the configuration and confirm it fails. A check that has never been seen to fail is an assumption.

Rollback is a plan, not a button

Teams get into trouble by treating rollback as a feature of the platform. It is not. It is a rehearsed sequence with dependencies, and if you have not walked through it you do not have one.

a laptop open on a desk beside a notebook, a phone and a cup of coffee

Things that commonly break the assumption:

  • The schema change is not backwards compatible, so rolling back the code leaves code that cannot read the data. This is why the contract step exists.
  • The rollback deployment itself fails, for reasons that range from an expired credential to a pipeline that assumes forward-only migrations.
  • The state changed between deploy and rollback. A background job wrote data in the new format that the old version cannot read.
  • Nobody knew who was authorised to make the call. During an incident, ambiguous ownership adds minutes.
  • The rollback path assumed a warm environment that has since been scaled to zero.

Rehearse it on a schedule rather than during an incident. Pick a recent release, roll it back in a non-production environment that resembles production, and time it. The exercise is worth doing for the discovery rather than the result, because you will find that one of your steps depends on a person, and that one of your pipeline stages has not worked since a credential rotation.

It also clarifies the decision rule, which should be written down before you need it. What error rate triggers a rollback, who decides, and how long the decision can take. Teams that have agreed this in advance roll back quickly. Teams deciding under pressure tend to wait too long, and by the time they act the original problem has compounded with the effects of the bad release.

Agreeing the risk appetite

Zero downtime is not free, and pretending otherwise produces estimates that cannot be met. The realistic costs are additional infrastructure during a blue and green window, the engineering time to make the schema backwards compatible, the monitoring needed to judge a canary, and the rehearsal time for rollback. None of these show up as a single line item, which is why they get forgotten.

a desk by a window with a laptop open in daylight

The conversation to have with a client early is about risk appetite rather than technology. Which journeys cannot be interrupted at all. Which can tolerate a short outage. How quickly a bad release must be detected and reversed. How much the organisation is willing to pay for infrastructure to hold a second environment. What happens during the periods nobody planned for, such as a change that lands just before a period-end freeze.

Those answers determine the approach, and they are the client’s to give. SmartEdge IT Solutions explains the trade-offs in terms they can decide on rather than in terms of deployment tooling. That work usually happens during application deployment planning, and the infrastructure decisions underneath it tend to belong to cloud infrastructure management or DevOps consulting depending on how much of the estate is already automated.

One final piece of honesty. Zero downtime covers the deployment, not the system. A release can complete with no user-visible interruption while a background job corrupts data, a queue silently drops messages, or a third party degrades. Plenty of organisations have discovered this the hard way, having invested in a careful release process and then lost an afternoon to a problem that no deployment strategy would have caught. Continuity of the release is worth doing properly. It is not the same property as reliability of the product, and conflating them is how the phrase gets oversold.

Editorial profile

Hannah Mitchell Infrastructure and Reliability Editor

Hannah Mitchell covers cloud, servers, deployment and running software reliably for SmartEdge IT Solutions. Her articles come out of operational reality: backups that were never tested, bills that hid real waste, and outages that had a cause somebody could have named.

Also 2 articles in the Insights archive.

← Back to Blog