Home / How we monitor

How we monitor the systems we manage

The point of a managed IT contract is that you find out about problems from us, not from a member of staff standing at your desk saying the system is down again.

That is easy to put on a website and harder to actually build. "We monitor your systems" is the vaguest sentence in our industry - it covers everything from a genuine 24/7 alerting chain down to somebody logging in on a Monday to see whether anything looks off. So this page sets out exactly what we watch, what reaches a human being, what does not, and where the honest limits are.

Four layers, and why you need all of them

Monitoring one layer well is worse than it sounds, because failures rarely announce themselves where you are looking. A server can be up, responding to pings, and completely useless because the application on it has stopped. A website can load perfectly while the backup that protects it has been silently failing for six weeks. We watch four separate things.

1. Is the hardware healthy?

Disk space, disk health, memory pressure, CPU saturation, temperature, RAID status, UPS battery condition. This is the unglamorous layer, and it is where most preventable outages get caught. Drives fail slowly and noisily before they fail suddenly and quietly. A disk filling at a predictable rate gives you two weeks of warning if anyone is reading the graph.

2. Is the service actually working?

A separate question from whether the machine is on. We check the things your staff and customers actually touch - the line-of-business application, the file share, the mail flow, the website, the VPN endpoint - from outside your network, on a schedule, the way a real user would reach them. Checking a service from the same machine it runs on is a comfortable illusion; it passes right up until the moment the whole box is the problem.

3. Has something changed that nobody authorised?

New administrator accounts, changes to firewall rules or Group Policy, unexpected modifications to system files and scheduled tasks, sign-ins from places nobody is. This is the security layer, and it is the one that separates a monitoring setup from a genuine security posture. Ransomware does not announce itself with a support ticket. It shows up first as small, boring changes that look like admin work.

4. Did the backup actually work?

Not "did the job report success" - whether the data can be restored. A backup nobody has ever restored from is a hope with a schedule attached. We verify restores periodically and time them, because the number that matters in a crisis is not whether you have a backup, it is how many hours until people can work again.

The test we apply: if a critical file changed on your server right now, when would you find out - and would the alert tell you which account did it? For most businesses arriving with no managed monitoring, the honest answers are "eventually" and "no".

Not everything is an alert

The failure mode of enthusiastic monitoring is noise. Turn everything on and route it all to email, and within a fortnight the notifications are wallpaper - people filter them into a folder they never open, and the one that mattered scrolls past at 3am with the other four hundred. An alerting system nobody reads is functionally identical to no alerting system, except it costs more and provides false comfort.

So most of what we collect is never sent to anybody. It goes into the record, it feeds the trend lines, and it is there when we need to answer a question after the fact. Only a small fraction is genuinely worth interrupting a human being for, and deciding which fraction is the actual skill. A patch window touching hundreds of files is not an incident. The same pattern of file changes on a Sunday afternoon, from an account that does not normally do that, is.

When something does fire

Alerts route to a named person with a timeout, not to a shared mailbox and a hope. If nobody acknowledges within the window, it escalates to the next person. Severity determines the channel: a disk at 85 per cent becomes a ticket and a scheduled fix, while a domain controller going offline becomes a phone call regardless of the hour.

Out of hours, the distinction we care about is whether the problem is costing you money right now or merely will on Monday. A failed backup job at 2am is important and can wait until morning. A production server that has stopped responding cannot, and that is what the after-hours path exists for.

We run this on our own systems too

It would be a poor look to sell monitoring and not use it. Our own web presence and internal systems sit under the same external checks we apply to client services, using AlertKick - continuous checks from outside our network, TLS certificate expiry tracked on the same schedule, and alerts that escalate to a person rather than into a mailbox. It is the same tooling and the same standard, which is the only honest basis on which to recommend either.

What monitoring does not do

Two things worth saying plainly, because the industry tends not to.

Monitoring is detection, not prevention. None of it stops a bad thing happening. It shortens the distance between the bad thing happening and somebody competent knowing about it - which is genuinely most of the value, because almost every expensive IT disaster is a small problem that had hours or days to grow while nobody was watching. But the prevention work is separate: patching, removing local administrator rights, retiring hardware before it ages into the danger zone, and keeping the number of people with domain admin small enough to name out loud.

It is not a substitute for backups. Monitoring will tell you precisely which account started encrypting the file share and at what time. It will not give you the file share back. That is what tested, off-site backups are for, and testing them is the part almost everyone skips until the week they cannot.

If you are not sure which of the four layers you currently have - and most businesses we assess have one and a half of them - that is exactly what the assessment is for. See our case studies for how this has played out for other businesses in the Lower Mainland, or read our take on what an hour of downtime actually costs.

Would you know before your staff did?

Book a free assessment and we will tell you exactly what is being watched today, what is not, and which gap is most likely to cost you first.

Book a free IT assessment