Platform

New Relic Implementation for Platform Monitoring

The team should hear it first.

New Relic is how a team sees what the platform is doing while it does it — response times, error rates, database calls and what a real visitor actually waited for. The value is not the dashboard. It is the alert that arrives before the phone call.

Where it stands
In production

New Relic is the monitoring platform with the whole picture in one place — application timing, real-user data from the browser, synthetic checks, logs and traces, and alerting that routes to a person. Most organizations buy monitoring after an outage and configure it in the mood the outage created, which produces a wall of alerts no one can act on and a dashboard no one opens twice. The discipline is subtraction. A useful monitoring setup answers three questions and refuses the rest — is the platform up, is it fast enough for the people using it right now, and is something degrading that will become an outage next week. Everything else is a chart. The judgment is knowing which numbers deserve to wake someone, which belong in a weekly review and which exist only because the tool collects them by default.

How the Work Splits

New Relic provides

Application performance monitoring across the stack, real-user monitoring from the browser, synthetic checks against critical paths, infrastructure and log data in one place, distributed tracing across services, and alerting with routing and escalation.

Pare & Co provides

The decisions the tool cannot make. What to instrument and what to leave alone, which thresholds mean something for this platform rather than for platforms in general, alert routing that reaches a person who can act at the hour it fires, and the discipline to delete an alert that has cried wolf twice. We also connect what monitoring sees to what the pipeline does, so a regression that shows up in production has a test that catches it next time.

Together

A platform whose team finds out before its audience does. Instrumentation on the paths that matter, thresholds set against how this platform actually behaves under its own traffic, and a short list of alerts that are worth interrupting someone for.

The work in practice

Monitoring is the part of a platform no one thinks about until the week they think about nothing else. It gets bought during an incident, configured in a hurry, and then quietly stops being read — because the thing that was installed answers a hundred questions and the team only ever had three.

The three questions worth instrumenting for

Is it up? The cheapest and least interesting of the three, and the one every tool does. A synthetic check against the paths that matter — the homepage, the search, the form that takes an application — beats a check against the root URL, because a site can return 200 with its most important journey broken.

Is it fast enough for the people using it now? Server-side timing is what the platform experienced. Real-user monitoring is what a person on a four-year-old phone on a rural connection experienced, and the two numbers separate more than teams expect. The second one is the number that affects conversion and search visibility, so it is the number worth alerting on.

Is something degrading? The outage that surprises a team is almost always the one that had been announcing itself for a fortnight — a slow query getting slower, a memory ceiling creeping up, an error rate at half a percent that used to be at a tenth. This is the question dashboards are actually good at, and it is a weekly-review question rather than an alert.

The failure mode is noise, not blindness

An alert that fires and does not need action teaches the team to ignore alerts. Two of those and the channel is dead — and the outage arrives in a channel everyone has learned to skip.

So the rule we work to is that every alert names a person and an action. If no one can say what they would do when it fires, it is a dashboard row, not an alert. And an alert that has been wrong twice gets deleted or given a new threshold rather than tolerated, because leaving it costs more than the thing it was watching.

Where it meets the pipeline

Monitoring that only reports is half a system. What production sees should change what the pipeline checks — a regression that reached real users is a regression with no test behind it, and the fix is a test as much as it is a patch. That loop is the reason this sits inside DevOps rather than beside it: instrumentation, thresholds, alert routing and the automated checks that run on every release are one practice, not two.

If you are weighing it

The question is rarely which tool. It is who gets woken, for what, and whether that person can do anything at that hour. Answer those three and the tool matters much less than the vendor comparison suggests. Answer none of them and the most capable monitoring product on the market will produce a very detailed record of an outage no one noticed.

Practice leadership

Need to know before your users do?

Tell us about it.
Start a conversation