Contact
Website Support & SLAs

Website Emergency Support: Retainer, Ad Hoc or None at All

This page compares what a website emergency support retainer actually buys against calling an agency cold mid-incident, and is written for IT managers, digital leads and marketing managers responsible for a site where downtime costs sales, leads or public trust.

The moment you need emergency support is the worst moment to procure it

Emergency website support is nearly always bought under duress. The site is returning 500s, the phone is ringing, and someone is opening a browser tab to search for an agency that can help right now. That is procurement at the worst possible moment: no time to check references, no time to negotiate terms, no ability to compare rates, and no leverage at all.

The uncomfortable part is that the vendor you reach in that state cannot help you quickly even if they want to. They have never seen your server. They do not know whether the site is on a single EC2 instance or behind an autoscaling group, whether the database is RDS or local, whether deploys go through a pipeline or over SFTP, or whether anyone has a current backup. Before they can diagnose anything, they have to discover everything. That discovery happens while your site is down.

Organisations often assume the choice is between paying for a retainer they may not use and paying hourly only when something breaks. That framing misses the real variable, which is not cost. It is elapsed time, and elapsed time is decided almost entirely by work done before the incident.

What actually slows down an emergency response with an unfamiliar vendor

In practice, the clock in a cold-start incident is consumed by administration, not engineering. The recurring blockers look like this:

  • Nobody can authorise spend. An unfamiliar vendor will not start billable emergency work without a purchase order, a signed scope or a credit card. Finance is not on shift at 9pm.
  • Nobody can grant access. Server SSH keys sit with a former developer. The cPanel login is in a departed marketing manager's password manager. Registrar access requires a security question nobody remembers.
  • The hosting account is not yours. The site is on a reseller account controlled by the original build agency, and their support desk opens at 9am.
  • No one knows what changed. Without deployment logs, change records or a plugin update history, the responder cannot correlate the failure with a cause and is reduced to guessing.
  • The backup position is unknown. The fastest safe action in many incidents is a restore. If nobody can confirm when the last verified backup ran and what it contains, restoring becomes a risk rather than a remedy.
  • There is no monitoring history. Without response time and error rate data from before the failure, there is nothing to compare against and no way to tell a gradual degradation from a sudden break.

None of that is engineering difficulty. It is preparation debt, and it is paid back in minutes of downtime at the least convenient moment. Our guide to the first hour of an urgent website incident covers what to gather while you are still deciding who to call.

Comparison: no arrangement, ad hoc hourly, retainer, full SLA

The four common arrangements differ far more in behaviour during an incident than they do on an invoice.

Attribute No arrangement Ad hoc hourly Emergency retainer Full managed SLA
Who answers Whoever replies to a cold enquiry A vendor you have used before, if available A named team with your account on file An on-call roster with a documented escalation path
Time to first human response Unpredictable, often next business day Hours, subject to their current workload Contracted, typically measured in minutes to hours by severity Contracted by severity, with after-hours and 24/7 tiers
Commercial approval needed mid-incident Yes, before any work starts Usually yes No, drawn from prepaid hours No, covered by the agreement
Access and credentials Collected during the outage Partially held, often stale Held and periodically verified Held, verified, plus infrastructure owned or co-managed
Environment knowledge None Whatever they remember from the last job Documented in a runbook Continuous, from monitoring, patching and deploys
Monitoring and detection You find out from a customer You find out from a customer Usually yours to provide Provider detects and often responds before you report it
Contractual remedy if they are slow None None Service credits or termination rights, if negotiated Defined credits, reporting and review obligations
Cost behaviour Highest hourly rate, uncapped Standard or emergency loaded rate Fixed monthly, drawdown against hours Fixed monthly, includes prevention as well as response

The pattern is straightforward. Each step up the table converts unknown elapsed time into known elapsed time. That is the product you are actually buying.

Response time versus resolution time, and why only one can be guaranteed

Every credible support agreement guarantees response time and no credible agreement guarantees resolution time. This is not an evasion, and you should be suspicious of any vendor who promises otherwise.

Response time is the interval between a ticket being raised, or an alert firing, and a qualified engineer acknowledging it and beginning work. It is entirely within the provider's control, because it depends on rostering, alert routing and staffing. It can be measured and it can be breached.

Resolution time depends on the fault. Restoring a corrupted database from a six-hourly snapshot is a bounded task. A volumetric DDoS, an upstream provider outage, a payment gateway API failure or a zero-day in a CMS core with no patch available are not bounded by anything the provider controls. A vendor who guarantees a fix in four hours is either excluding everything difficult in the fine print or intends to miss.

What a mature agreement adds instead of a resolution guarantee is update cadence: a commitment to communicate status at defined intervals while an incident is open, for example every 30 minutes for a critical outage and every business day for a low-severity defect. That is the thing stakeholders actually need, because the second-worst experience after an outage is silence during one. If your current provider is failing on either measure, the practical steps are set out in our note on what to do after a website SLA breach.

Severity definitions that hold up when everyone disagrees about severity

Severity arguments happen because the client rates by frustration and the vendor rates by effort. Both are wrong. Severity should be defined by business impact and written down before anything breaks, so that classification is a lookup rather than a negotiation.

Severity Definition Examples Typical target response
P1 Critical Site or a revenue-critical function is unavailable, or a confirmed security compromise is in progress Site returns 5xx, checkout fails for all users, defacement, active data exfiltration Minutes, 24/7
P2 High Major function degraded with no workaround, or severe performance loss Search broken, forms not delivering, page loads exceeding several seconds sitewide Within business hours, same day, or extended hours by agreement
P3 Medium Function impaired with an acceptable workaround One template rendering incorrectly, a broken integration with a manual fallback Next business day
P4 Low Cosmetic, content or enhancement requests Typographic issues, minor CSS defects, small content changes Scheduled into the maintenance queue

Two clauses make this workable. First, the client can raise severity once without justification, which prevents the vendor from quietly downgrading inconvenient tickets. Second, severity is reassessed as facts emerge, so a suspected P1 that turns out to be a single user's caching problem is re-rated with a note rather than an argument. Full worked examples of these definitions appear in our page on website support SLAs for enterprise and government.

Access, credentials and environment knowledge as prerequisites for speed

A response time commitment is only honest if the responder can actually act. Before any agreement starts, the following should exist and be verified, not assumed:

  • SSH or console access to every server, held in a shared vault with named individuals and MFA enforced.
  • Registrar and DNS control, tested by making a low-risk change, because a DNS emergency at 2am is not the time to discover the registrar login belongs to someone who left in 2019.
  • CDN and WAF administration, so rules can be changed under attack rather than filed as a support request with a third party.
  • CMS administrator accounts for each environment, plus database credentials and connection details.
  • A documented deployment path: repository, branch strategy, build pipeline and rollback procedure, with the rollback actually exercised at least once.
  • Backup inventory: what is backed up, how often, where it is stored, how long it is retained and when a restore was last tested end to end.
  • A runbook naming the stack, PHP or Node version, dependencies, cron jobs, integrations and any known fragile areas.
  • An escalation list with client-side decision makers, including who can authorise taking the site into maintenance mode.

Assembling this is usually a day or two of work. Doing it during an outage adds hours to every incident for the life of the relationship, which is why we start engagements with an infrastructure audit rather than a login handover.

Business hours, after hours and 24/7 cover: what each realistically means

Cover windows are frequently misread. Business hours cover means an engineer is at a desk and alerts are watched between, say, 8am and 6pm AEST on business days. An incident at 7pm Friday is picked up Monday morning. That is entirely reasonable for a brochure site and entirely unreasonable for a retailer in November.

After-hours cover means someone carries a pager outside business hours and responds to alerts above an agreed severity. Response is slower than during the day, because a person has to wake up, get to a laptop and load context. A target of 30 to 60 minutes overnight for a P1 is realistic. Fifteen minutes, overnight, from a small team, usually is not.

True 24/7 cover requires a roster, with primary and secondary on-call, handover between people and time zones, and alerting that routes to a human rather than an inbox. A distributed team across Australian and Asian time zones makes this considerably more sustainable than asking one Sydney engineer to be permanently awake.

Whatever window you buy, it is worth nothing if detection is slow. If you learn about outages from customers, your effective response time is however long it takes someone to complain plus however long the vendor takes to answer. Pair any cover window with availability monitoring that combines synthetic checks with real user data so the clock starts at the failure, not at the phone call.

Post-incident review and stopping the same outage twice

Ad hoc emergency work has a structural flaw: the vendor is paid to make the symptom go away and is not paid to prevent it recurring. Restarting PHP-FPM clears a memory exhaustion event. It does not address the unbounded query that caused it, so the same call happens again in a fortnight, at the same hourly rate.

A retainer or SLA should include a written post-incident review for every P1 and significant P2, produced within a few business days, covering:

  1. A timeline from first symptom to resolution, with timestamps for detection, acknowledgement, diagnosis and fix.
  2. The direct technical cause, stated plainly.
  3. The contributing conditions, for example missing alerting thresholds, an unpatched dependency or an undocumented manual change.
  4. Remedial actions with owners and dates, separated into immediate fixes and structural changes.
  5. What detection would have caught it earlier, and whether that check is now in place.

Over a year of these, the character of the relationship changes. Emergencies become rarer because the causes are being removed rather than reset. That is the strongest financial argument for a retainer, and it is invisible on a per-incident invoice. It is also the practical difference between break/fix emergency support and a managed arrangement where prevention and response come from the same team.

Choosing between the three arrangements

Ad hoc hourly is defensible when the site is not commercially critical, you hold all credentials, and an outage lasting a day is an inconvenience rather than a loss. It is a genuine option and not everyone needs more.

An emergency retainer suits organisations with an in-house team that handles routine work but has no depth at 11pm, and no one who has restored a production database under pressure. You are buying a phone number that answers and a team that already knows the stack.

A full managed SLA suits organisations where the website drives sales or leads, where government or enterprise procurement requires documented response commitments, and where there is currently a gap between the hosting provider who will not touch the application and the build agency who has moved on. When the infrastructure and the application expertise sit with one accountable party, there is nobody left to point at.

Frequently Asked Questions

What is website emergency support?

Website emergency support is a service that responds to urgent faults such as outages, security compromises, failed deployments and severe performance degradation, usually outside normal project work and often outside business hours. It can be purchased ad hoc at an hourly rate, as a prepaid retainer, or as part of a managed service agreement. The critical difference between the models is how quickly a qualified engineer can begin work, which depends far more on prior access and environment knowledge than on hourly rate.

What is the difference between break/fix support and a website support SLA?

Break/fix support is transactional: you report a problem, the vendor quotes or bills hourly, and the engagement ends when the symptom is gone. A website support SLA is a continuing agreement with defined severity levels, contracted response times, agreed cover hours and remedies if those commitments are missed. Break/fix is priced per incident and rarely addresses root causes, while an SLA includes monitoring, prevention and post-incident review.

Should an SLA guarantee how long it takes to fix my website?

No. Response time can be guaranteed because it is within the provider's control, but resolution time depends on the nature of the fault, third-party systems and factors such as upstream outages or unpatched vulnerabilities. A credible agreement guarantees acknowledgement within a defined window and commits to a communication cadence while the incident stays open. Treat any promised fix time with scepticism and read the exclusions carefully.

How quickly can a new vendor help if my website is down right now?

An unfamiliar vendor typically needs commercial approval, server and CMS credentials, DNS and registrar access, and an understanding of the deployment and backup position before any diagnosis begins. That preparation commonly takes longer than the technical fix itself. If you are in an incident now, gather hosting details, recent change history and the last known-good backup date before you make the call, because it materially shortens the response.

Is a support retainer worth it for a site that rarely breaks?

It depends on what an hour of downtime costs you, not on how often incidents occur. If the site generates sales or leads, or serves a public obligation, the retainer is insurance against the unbounded cost of a cold-start response at the worst time. If the site is informational and a day offline is tolerable, ad hoc hourly support is a reasonable choice.

Can emergency support cover a website someone else built?

Yes, and it is a common starting point. It begins with an infrastructure and code audit to document the stack, verify access, test the backup and restore path and identify the fragile areas, because response commitments cannot honestly be made against an environment nobody has examined. Once that baseline exists, severity definitions and response times can be agreed with confidence.

Put emergency website cover in place before you need it

UnDigital provides managed hosting and website management under defined support SLAs for enterprise, government and growing businesses, including after-hours and 24/7 cover with documented severity levels, response targets and post-incident review. Because we hold both the infrastructure and the development capability, there is no handover between a host who will not touch the application and a developer who cannot reach the server. Engagements start with an infrastructure audit so the response commitments we make are ones we can keep. Explore our managed hosting for enterprise websites or get in touch to scope cover for your site.

Get a response time in writing

Reviews from our client partners.

"Thanks so much for your comprehensive strategy and execution of our digital ecosystem.

I can finally sleep at night knowing that everything is under control, secure and scalable.

Thank you!!!".

Corporate Marketing Manager, Sekisui House

"Thanks for all your help. This project was in such good hands from the beginning. We really appreciate all your hard work and expertise!!"

Retail Marketing Manager, West Village

Scope emergency cover for your site

@undigital