Service

You should hear it from us, not from a client.

Uptime checks every five minutes on the things a business actually depends on, with alerts that arrive when something changes state — and carry the fix with them.

The problem

Most outages are found by the person they inconvenience.

The usual way a small business learns that something is down is that somebody cannot do their job — or worse, a client mentions it first. By then it has been broken for an unknown length of time, and the first job is working out how long.

The opposite failure is noise. Monitoring that emails every five minutes while something is down gets muted inside a week, and once it is muted it may as well not exist.

Useful monitoring is quiet. It tells you when the state changes, once, and it tells you enough to act on without opening a laptop.

What’s included

What the work actually covers.

Checked every five minutes

The services that matter are polled on a five-minute cycle, so the gap between something failing and somebody knowing is measured in minutes rather than in however long it takes a person to bump into it.

Alerts on change, not on repeat

An alert fires when something goes from working to not working, and again when it recovers. It does not repeat every cycle while the problem lasts, because that is precisely the behaviour that trains people to ignore alerts.

The fix travels with the alert

The message carries the recovery step for that particular service, so the first response does not depend on finding documentation, a laptop, or the office.

Brief blips stay quiet

A service that fails one check and passes the next does not produce a pair of alerts. The state has to settle before anything is sent, so a momentary wobble does not wake anybody at two in the morning.

It reaches a phone

Alerts arrive as a push notification rather than an email sitting in a queue nobody is watching outside working hours.

Written down like everything else

What is monitored, how often, what counts as a failure and who gets told — recorded and handed over, rather than living only in our head.

How it runs

From nothing to notified.

Decide what actually matters

Not everything deserves an alert. We pick the handful of services where being down genuinely stops work, because monitoring everything is the fastest route to ignoring all of it.

Check the right thing

Each service gets a check that proves it is really answering, not merely that a machine somewhere replies to a ping. Those are very different questions.

Tune the noise down

Thresholds and settle times get adjusted over the first few weeks, until alerts are rare enough that receiving one means something.

Hand it over

You get the list of what is watched, how, and where the alerts go. Replace us and the monitoring keeps running, and somebody else can change it.

Questions

Asked before every job.

Is this full network monitoring?

No, and it is worth being straight about it. This is uptime and service monitoring — is the thing answering, and if not, when did it stop. Polling switches and interfaces over SNMP is a larger build and a separate conversation.

Am I going to be woken at three in the morning?

Only if you want to be. Which services can alert out of hours, and who they reach, is a decision we make deliberately rather than something you inherit by default.

What does it run on?

Infrastructure we operate, so there is nothing extra for you to host or patch. If you would rather it ran on something of your own, that works too.

How many alerts are normal?

Very few. If alerts are arriving weekly, either something is genuinely wrong or the tuning is, and both are worth fixing rather than tolerating.

Tell us what you would want to know about first.

Name the one system where being down for an hour would actually hurt. The first conversation is free.