This website uses cookies

Read our Privacy policy and Terms of use for more information.

Vendor-neutral · Weekly since Feb 2024

Own the signal.Rent the platform.

Observability strategy, incident leadership and AI operations for the people who are accountable when it breaks. Written by Allan Mann, with nothing to sell you but the thinking.

One email a week · Unsubscribe in one click
checkout · p99 latency ◆ breach 14:22
SLO 400ms -6h to now
Everyone sees the spike. The argument is about who owns the fix.
561
Subscribers
36.2%
Open rate
147
Pieces published
103
Issues of The Signal
Who writes this

Allan Mann

Twenty-five years running IT operations and observability across five tier-1 banks and one UAE government, the environments where "it looked fine on the dashboard" ends up in front of a regulator.

Author of Metrics & Mayhem: A CTO's Guide to Observability That Actually Works. Host of the Metrics & Mayhem podcast. Vendor-neutral by construction: no platform reselling, no vendor sponsorship, no referral fees.

PRA FCA Bank of England Vendor-neutral
A global tier-1 bank Modernised monitoring across 40+ countries.
A UAE government platform Built the observability function from nothing, for a platform serving 3.2 million citizens.
A tier-1 UK bank Designed metrics and the Dynatrace business-events programme for three of the largest payment applications, inside a payments-resilience initiative under PRA, FCA and Bank of England oversight.
Why this exists

Most observability advice is written to sell a platform.

This is not. No vendor sponsorship, no reference architecture that happens to need three seats of something. Just what actually holds up in a large enterprise, and what quietly does not.

The rollout

It was never the tooling

OpenTelemetry is not what is blocking you. Ownership boundaries, funding models and the team that will not go first are. Those are the pieces nobody writes about.

The bill

Cost is a design decision

Cardinality, retention and sampling get chosen by accident, then defended for two years. Decide them on purpose and the invoice stops being a surprise.

The agents

AI needs a gate, not a leash

Autonomous SRE agents will run queries you would never approve and spend budget while they think. The question is where accountability actually sits.

Free chapter

What separates the teams that recover fast from those that do not

Chapter 4 of Metrics & Mayhem. The chapter that answers the question every engineering leader is afraid to ask out loud: when it breaks, who actually owns getting it back?

  • Why one airline recovered from the CrowdStrike outage in a day and another took five, from the identical software fault.
  • The three-layer ownership model: outcome owner, indicator owner, platform owner.
  • Why most organisations build only the platform layer, then wonder why their P1s take an hour.
  • One concrete Monday action you can take before you close the page.
Metrics and Mayhem book cover

Recent writing

The thinking, in public

Where to find it

Three ways in

The Signal

One email every Friday since February 2024, issue 103 and counting. One idea, argued properly, for people who have to make the call on Monday.

Subscribe →

Metrics & Mayhem

The podcast. Short audio notes from the front line of IT operations, with deep dives when a topic earns the longer cut.

Listen →

The book

A CTO's Guide to Observability That Actually Works. Eleven chapters. Chapter 4 is free, and it stands on its own.

Read chapter 4 free →
Advisory

More tools. More dashboards. Worse recovery.

In 2021, 47% of organisations took more than an hour to recover from an incident. By 2024, it was 82%. Spend went up. Outcomes went down. Not a tooling problem. A leadership problem wearing a tooling costume.
Start here from £5,000 Fixed fee. Around five days. Bounded scope, agreed in writing up front.

The Observability Assessment

A fixed-scope review of one part of your observability: one platform, one service domain, or a defined set of questions. Not the whole estate.

  1. 01
    Prework, before I arrive
    An intake pack: access, current dashboards and alerting, recent incident data, and the people I need to speak to. The work does not start until this is in.
  2. 02
    Review and interviews
    I work through what you have supplied, and I talk to your people. The conversations matter as much as the data.
  3. 03
    Findings and roadmap
    A written assessment of what is working, what is theatre, and a prioritised roadmap, plus a readout session with your team.

Ongoing advisory from £1,250 per day, scoped and quoted per engagement. No retainer you have to justify.

One email a week.
No vendor pitch at the end of it.

Join 561 IT and engineering leaders who get The Signal every Friday.

Unsubscribe in one click · No sharing, ever