Services / Care / A real response promise

A response time you can hold us to, in writing

Every supplier describes themselves as responsive right up until something breaks on a Friday afternoon. A service level agreement turns that into numbers: how fast we reply, how fast we restore service, who you call, and what it costs us when we miss.

Why it matters

The cost is the wait

Ask a business owner what went wrong with their last technical supplier and you rarely hear about code quality. You hear about the message that went unanswered for four days, the ticket that vanished into a portal, the developer who read WhatsApp when it suited them, and the discovery, mid-crisis, that there was nobody above the person ignoring you. The technical work was probably fine. The accountability was not there at all.

An SLA fixes that by settling the arguments in advance. It defines severity levels, so nobody debates whether a broken checkout is urgent while it is broken. It sets coverage hours that match when your business actually trades. It separates response, which is entirely within our control and can be promised outright from repair, which depends on what broke, and which we commit to as a restore-or-workaround target rather than a fictional guarantee that every conceivable bug is fixed within four hours.

It also names people. A single front door for raising something, a named engineer who already knows your systems rather than whoever is next in a queue, an escalation path with actual names and numbers on it, and service credits when we miss. And because a promise nobody measures is just marketing, we report our own performance against it every month, including the months we did not meet it.

What you actually get

Built to be trusted

What an SLA with us actually contains, in the terms you would want it argued in.

01

Severity defined before it matters

A written matrix agreed with you: what counts as critical, major or routine, expressed in terms of business impact rather than technical symptoms. It means the middle of an incident is spent fixing things rather than negotiating how serious it is.

02

A response time in the contract

How quickly a human who can actually help acknowledges you, by severity and within your coverage hours. Not an auto-reply confirming receipt: a person, with your systems in front of them, telling you what is happening.

03

Restore first, fix properly after

The first obligation is getting you working again, which often means a rollback or a workaround rather than the perfect solution. The root cause is then dealt with in daylight, with a written explanation of what happened and what stops it recurring.

04

Coverage that matches when you trade

Business hours, extended hours, weekend cover for the retail peak, or genuine round-the-clock. Priced accordingly, and chosen from what your business actually does, buying overnight cover for a system nobody touches at night is money set on fire.

05

A named engineer, not a rotating queue

Someone who knows your architecture, your quirks and your history, with a documented second who also knows it. Continuity is what makes a fast response possible; explaining your system from scratch every time is why other arrangements feel slow.

06

We measure ourselves and show the misses

Every ticket timestamped against its target, reported monthly, with service credits when we fall short. Publishing your own failures is uncomfortable, which is exactly why it is the part that makes the rest credible.

Where it earns its keep

Same pattern, different desks

What an SLA needs to say depends entirely on what an hour of downtime costs you, which is a business question, not a technical one.

Retail & e-commerce

01 · Retail & e-commerce

The weekend is not a quiet period

The problem
Your heaviest trading happens on evenings, weekends and the run-up to Christmas, which is precisely when standard business-hours support does not answer. A checkout fault at 6pm on a Saturday in December is not an inconvenience. It is a substantial share of the quarter.
What we build
Coverage hours built around your actual trading pattern rather than an office calendar, with an enhanced peak-season tier, a change freeze through the busiest weeks, and severity definitions written in commercial terms. Anything blocking payment is critical by definition, regardless of what the technical cause turns out to be.
What changes
The busiest hours are the best covered rather than the worst, and there is a named person reachable during them with the authority to act.
Healthcare & clinics

02 · Healthcare & clinics

When the booking system is the front door

The problem
Patients book, reschedule and complete forms online, and the clinic runs its day from the same system. An outage during opening hours is chaos at reception within minutes, and there are records obligations that do not pause because a supplier is unavailable.
What we build
Tight response targets during clinic hours with a lighter overnight tier, a documented manual fallback so reception can keep working while a fix is applied, severity levels that treat anything touching patient data as critical, and an escalation path that reaches a decision-maker on both sides quickly.
What changes
The practice has a rehearsed plan rather than an improvised morning, and the obligations around access and notification are handled by people who already know the deadlines.
Manufacturing & distribution

03 · Manufacturing & distribution

When software stopping stops physical work

The problem
The warehouse system tells pickers what to pick and prints the labels. When it stops, people stand still and the cost per hour is a number you can calculate exactly, which makes it very obvious how inadequate a next-business-day arrangement is.
What we build
Severity defined by operational impact, with the picking and despatch paths as their own top tier; response targets aligned to shift patterns including early starts; an agreed offline procedure so a shift can continue degraded; and quarterly rehearsal of the failover so the procedure is something people have practised.
What changes
An outage becomes a slower shift rather than a stopped one, and the SLA is sized against a downtime cost everyone has agreed rather than guessed at.

The technology

The tools behind it, named

An SLA is a contract, not a product, but it is only deliverable if the machinery behind it exists. This is what has to be in place before a number in a contract means anything.

6 layers · 25 technologies

01

Where a request lands

One front door, timestamped, so nothing depends on someone remembering a conversation. You use whichever of these your team already lives in.

  • Linear
  • Jira
  • Zendesk
  • Intercom
  • Slack

02

How the clock starts without you

Most incidents should open themselves. Monitoring raises the ticket and starts the response clock before anyone at your end has noticed.

  • Sentry
  • Grafana
  • Uptime Kuma
  • Better Stack

03

Getting hold of a human

On-call rotas, acknowledgement tracking and automatic escalation when the first person does not pick up, including out to a phone when a message will not do.

  • PagerDuty
  • Opsgenie
  • WhatsApp
  • Phone escalation

04

Saying so publicly

Handling an incident badly damages trust more than the outage does. A status page and a prepared notification pattern make honest communication the default.

  • Statuspage
  • Instatus
  • Resend

05

The runbook the fix comes from

Documented procedures, architecture diagrams and account registers kept current, so a response at 2am is a competent one rather than an investigation.

  • Notion
  • Confluence
  • GitHub
  • Mermaid

06

Restoring service quickly

A promised restore time is only credible if rolling back is a single reliable step that has been practised, not a heroic effort.

  • GitHub Actions
  • Docker
  • Cloudflare
  • Vercel
  • Terraform

Product names and logos are the property of their respective owners and are shown to describe the technologies we work with. Their use does not imply any partnership, sponsorship or endorsement.

How we deliver it

Live behind a human first

Agreeing an SLA takes a couple of conversations. Being genuinely able to meet one takes a little longer, and we will not sign numbers we have not tested.

01

We work out what an hour of downtime costs

Revenue per trading hour, staff standing idle, contractual exposure, reputational damage. It is a rough figure and that is fine. Its job is to tell us whether you should be buying overnight cover or whether that would be waste.

02

We write the severity matrix together

Each level defined by business impact with concrete examples from your own system, so both sides recognise a critical issue the same way. This single document prevents most of the arguments that sour support relationships.

03

We agree the numbers

Response and restore targets per severity, coverage hours, escalation contacts on both sides, exclusions stated plainly, and the service credits that apply when we miss. Written into the contract rather than described on a web page.

04

We make sure we can actually meet them

Monitoring in place, access provisioned, backups restored as a test, rollback proven, runbooks written. You cannot honestly promise a restore time you have never measured, so this step happens before the agreement starts rather than after.

05

We set up the front door

One channel for raising things, an on-call rota with a named primary and second, automatic escalation on non-acknowledgement, and a short briefing so your team knows exactly what to do and who to call at 8pm.

06

We report against it, then review it

Monthly figures on tickets, response times, misses and their causes. Quarterly we revisit whether the levels still fit: businesses change, and an SLA written for last year's traffic is often either too expensive or no longer enough.

Before you commit

The questions worth asking

What is the difference between a response time and a fix time?

Response is entirely within our control, so we promise it outright: a human who can help acknowledges you within the agreed window and tells you what is happening. Repair is not, because it depends on what broke. A third-party outage or a subtle data corruption cannot honestly be given a guaranteed clock. What we commit to instead is a restore-or-workaround target, continuous effort at the agreed severity until service is back, and regular updates throughout. Anyone promising a guaranteed fix time for every possible fault is describing a contract they intend to argue about later.

Do you offer round-the-clock cover?

Yes, and it costs meaningfully more, because it means a rota of people whose evenings and weekends are constrained. Most businesses we talk to do not need it, extended hours covering their real trading pattern, with an emergency escalation for genuine catastrophes, gives nearly all of the benefit at a fraction of the cost. We will tell you which one your business actually justifies.

What happens if you miss a target?

Service credits against the following month's fee, at the level set out in the agreement, applied automatically rather than on request. You should not have to chase a discount you are owed. You also get a written post-incident review covering what happened, why we missed, and what changes so it does not repeat. Credits are not really the point; they exist to make sure the report is honest.

Is the retainer a blank cheque for new features?

No, and it is worth being blunt about that because it is the most common misunderstanding. The SLA covers keeping what exists working: incidents, faults, patching, monitoring. Improvement work is a separate defined allowance of hours, and anything substantial is scoped and quoted like the project it is. Suppliers who blur that line either pad the retainer heavily or start quietly deprioritising your support tickets, and neither ends well.

What if the outage is our hosting provider's fault, or a third party's?

We cannot fix another company's platform, and the agreement says so plainly. But the clock does not stop on the parts that remain ours: we confirm the cause quickly rather than making you wonder, put a workaround in place where one exists, keep you and your customers informed, and manage the supplier relationship on your behalf. Where a dependency fails often enough to matter, we will tell you what removing that single point of failure would cost.

We already have an in-house developer. Can you cover only out-of-hours?

Yes, and it is a sensible arrangement. Your team handles the working day; we take evenings, weekends, holidays and cover for leave. It needs shared runbooks and access on both sides, plus a clear handover rule for anything still open at the boundary, but it removes the single-person dependency that makes most in-house arrangements fragile.

What happens today if it breaks at six on a Friday?

If you do not have a confident answer, that is the thing worth fixing first. Tell us what you run and when you trade, and we will propose response levels sized to what downtime actually costs you.

Start the conversation