Why We Recommend Maintenance After a Build
There is an obvious commercial awkwardness in this article: we are recommending a service we sell. So the useful version is to be precise about which parts are engineering rather than sales, and to include a blunt account of what maintenance does not do.
The short version of the argument: a build is a snapshot. Everything that keeps software working stops working over time for reasons that have nothing to do with your code. Dependencies publish breaking changes, certificates expire, hosting providers change defaults quietly, and the one person who knew why something was configured that way leaves for a job that pays better.
What “support” should actually mean
Ambiguity here is the main cause of bad retainers, because “support” gets used for at least four different things while contracts buy one and expect another.

- Break and fix. Something is down or visibly wrong, and it needs restoring. Reactive, time-sensitive.
- Maintenance. Keeping the current version running as the world around it moves. Mostly scheduled, rarely dramatic, easy to postpone and expensive to postpone.
- Security upkeep. Patching, access review, certificate management, checking what changed.
- Change. Small new work: a new field, a new report, a different export. This is development, not support, and it is where retainer disputes come from.
Before signing anything, we like these defined in writing. Which of the four you are buying. Which hours are covered and what happens at the weekend. What a response time means in practice — “we begin investigating within four working hours” is a commitment somebody can keep; “we take it seriously” is not. Whether third-party systems are in scope, and if not, whether you will be asked to chase them. And what counts as a new project, so that the boundary does not get argued about in month six.
The failures we actually see
Not the dramatic ones. In order of how often they turn up:
- A framework or library publishes a version that changes a convention, and a deployment produces a blank page.
- A certificate expires. A domain, an API client, a mail server. Automatic renewal was never configured.
- The hosting provider changes something. A default runtime version, a minimum TLS setting, a mail relay, a disk performance limit on a smaller plan.
- A database major version upgrade goes through and a query plan changes, so a report that took two seconds now takes a minute.
- A third party tightens a rate limit or retires an endpoint, and a background job that used to take minutes now takes all night.
- Disk fills. Uploads, logs, or one runaway job nobody noticed. The site goes read-only and looks like a hack.
- Credentials expire: the SMTP password, a database user, an API token whose only copy is in a production environment variable.
- Nobody has run the application in months. The first deploy fails for reasons unrelated to the change being deployed.
- A hosting or vendor free tier ends without anybody being told, because it was on somebody’s card.
None of these is anybody’s fault on the day. They are the normal consequences of running software for a while, which is the entire argument.
Dependency updates are scheduled work, not luck
The temptation is to pin every version forever and never touch it. That trades a known, plannable problem for an unknown, unplannable one, and it also leaves the app on a version that eventually stops receiving fixes at all.

What works in practice is unglamorous and repetitive:
- Dependencies declared in a committed file, so the exact set is reviewable rather than implied.
- Update notifications turned on and routed to an inbox somebody actually owns.
- A monthly batch of patch-level updates, applied in staging and tested before production.
- A quarterly look at whether a major version is worth a planned piece of work, costed as its own item.
- Deployment kept small and reversible, so an update is a normal low-risk release rather than an event.
The reason this matters more than it looks: a framework upgrade in a repository with no meaningful tests is not an upgrade, it is a rewrite with a release note. That is also where CI/CD and deployment stops being a productivity story and becomes a risk story.
Monitoring that reaches you before the customer does
Monitoring nobody acts on is theatre. It looks responsible on a slide and provides no protection whatsoever, because the alert arrives in an inbox that a person reads on a good day.

So we are more interested in what happens after the alert than in how many things are monitored. A named route to a human. An escalation path when the first person does not respond. An explicit decision about overnight and weekend behaviour, including whether the site is degraded or left alone. And a short list of things that can be done immediately without waiting for approval — a restart, a rollback, a failover to a cached page.
What is worth watching is narrower than most checklists suggest. The user-facing flows that would be a business problem if they broke: form submission, checkout, enquiry delivery, login. Plus certificate expiry, disk and memory headroom, whether scheduled jobs actually succeeded, application error rates, and the health of third-party APIs you depend on. Monitoring and backup only earns its cost when the alerting path has been tested by something going wrong on purpose.
Backups are only real once they have been restored
A backup that has never been restored is a hypothesis. We have seen projects where the client was confident about recovery and the restore failed on the first attempt for reasons nobody had anticipated — an encrypted volume with a key that had been rotated, a dump that assumed an extension that was not installed, a file permission that made the copy unreadable.
Three properties make a backup real. It runs automatically, because manual backups stop at exactly the moment they are most needed. It is not stored only on the same machine or account as the thing it protects. And it is restored on a schedule into a scratch environment, by somebody who writes down how long it took.
That recorded restore time is the number that matters, and it is usually worse than the backup interval implies. Two daily backups a week apart can still mean losing a day of work. Three different terms get confused here: the recovery point is how far back you can go, the recovery time is how long restoring takes, and the backup frequency bounds the first. Which one your business actually needs is a decision, and it should be written down rather than assumed from a storage screenshot.
Incidents want a runbook, not improvisation
The most valuable document in most projects is a short runbook written before anything goes wrong: the page that decides, at two in the morning, who does what. It exists because during an incident nobody is inventing steps — they are looking things up.

A useful runbook covers the handful of scenarios that would genuinely hurt, with exact steps rather than principles. Who to contact, in what order, including the person outside the company. What may be done immediately without seeking approval, and what may not. How to tell customers if the site is visibly degraded, and who is allowed to send that message. And a template for the incident log, because reconstructing a timeline from memory afterwards is painful and inaccurate.
One capability above all others changes how an incident feels: being able to roll back to the previous release quickly. With that in place, most failures become an inconvenience with a small customer-visible window. Without it, every bad deploy is an investigation.
Being honest about the limits: a runbook written after an incident describes the incident you had, not the next one. We review them periodically and add scenarios rather than rewriting, because the pattern of near misses is a better guide than anyone’s imagination.
Small changes cost more than they look
A request described as “just a small change” arrives with a hidden fixed cost. Somebody has to read enough of the codebase to find where it belongs, work out whether it affects anything else, write it, test it, review it, deploy it, verify it in production, and update whatever documentation mentions it. Most of that overhead is independent of the size of the change.

This is why context transfer matters more than any individual handover session. Repository access for the people who will own it. An environment the client can actually run and break. Documentation that describes how to release, not just how the code is arranged. And names of the people who can answer questions when nobody remembers.
The other side of the same problem is change debt. Features get added without the tests, documentation or naming being updated, so the next change costs more than the last. After a few years the reason a small change is expensive is that nobody wanted to pay for tidying up at the time. Our testing and quality assurance work is usually partly about that, and partly about making the monitoring in the previous section worth having.
Retainer, on demand, or nothing
Three reasonable positions, and pretending otherwise helps nobody.
A monthly retainer suits an organisation with steady needs and real usage. Predictable cost, known capacity, someone who already knows the system, and no procurement queue at the moment a problem appears. The trap is buying a large one for occasional needs, then resenting it.
Pay as you go suits teams with a technically capable person internally who can triage and who can get approval quickly. Nothing is wasted, but you are buying response time you have not reserved, and it depends entirely on that internal person having both the skill and the availability.
Nothing at all is defensible more often than providers admit. A small site, low consequences, a modest number of changes and somebody internal who is comfortable with technology — that is a reasonable place to be. It stops being reasonable as soon as the system handles money, personal data or orders, or when the only person who understands it is on holiday.
The fourth option is hiring internally, which is right when the need is sustained rather than occasional and when somebody will manage the person rather than merely use them. The failure mode there is a lone owner with no peer review, which is how a support arrangement becomes a dependency. We are happy to talk about how hiring developers from SmartEdge IT Solutions works when that is the shape of the requirement.
What is not in a maintenance retainer
Worth being blunt about, because this is where the disappointment comes from. A maintenance retainer does not include new features or a redesign. It does not include migrating you to a different platform — that is a project, and pretending otherwise produces a retainer nobody wants. It does not include third-party fees, content production or licences of any kind. Emergency call-out outside agreed hours is usually chargeable, and we would rather agree that in advance than have it appear on an invoice.

Exit terms deserve more attention than they usually get. Who owns the source code (you, in every sensible arrangement). What documentation and infrastructure diagrams you receive. How credentials are handed back and transferred. What happens to monitoring accounts, domains and third-party subscriptions, which are usually registered to the provider’s address and then take a fortnight to move. Those details are far less painful agreed at the start than during a disagreement.
And if the honest answer is that you need less maintenance than you think you do, that is a perfectly good outcome. Several times the recommendation we have made at SmartEdge IT Solutions was a smaller arrangement than the one requested, or a fixed piece of work followed by nothing at all. Our application maintenance and support page is worth paying for when something real would break without it, and not before.
