We embed with your teams. We sit in on incident reviews, GameDays, chaos experiments, and operational readiness reviews, and we interview people at every level, from the engineers on call to the CEO.

The diagnosis names the patterns behind your recurring incidents. Most of them sit outside the engineering practices: how leaders behave under pressure, what gets rewarded, and how software and data are built and operated.

The diagnosis is where the work starts. We work with your teams for six to twelve months on the conditions it found. When an incident happens during the engagement, we analyse it together as it happens. Every practice we build is designed so your team runs it after we leave.

We ask what conditions made the failure possible, and why the decisions involved made sense to the people who made them at the time. We look for what surprised people, and for the weaknesses that haven't caused an incident yet. The analysis holds multiple teams accountable without blaming any of them, and puts the gaps in the system rather than in individuals.

No prescriptive checklists. We study your organisation's culture first

Founder-led. The person you talk to is the person doing the work

Diagnosis before prescription. We don't sell solutions to problems we haven't seen

What you need to hear, not what you want to hear

Every organisation has a gap between what it thinks happens in its systems and what actually happens. The gap reopens with every change. Learning is how the organisation closes it, and it happens in three steps.

Sense the gap: incidents do it on production's schedule; load testing, chaos engineering, and GameDays do it on yours. Understand it: incident analysis after the fact, operational readiness reviews before the next change. Improve: change the system, the practices, and what people believe. Then sense again. Resilience is closing the gap at least as fast as it reopens, and before production does it for you.

Sense, understand, improve The gap between what you think happens and what actually happens sits in the centre. Around it, a cycle of three steps: sense the gap through incidents or experiments, understand it through analysis and reviews, improve the system and the team, then sense again. THE GAP What you think happens vs what actually happens reopens with every change Sense the gap incidents, experiments Understand it analysis and reviews Improve the system and the team

Every practice can run in two modes: as a ritual that performs diligence, or as a mechanism that feeds what it finds back into the loop. Which mode you get depends on what sits under every practice: psychological safety, the right incentives, and leadership support. That is the bedrock. Without it, no practice works.

The assessment finds where your loop is broken: gaps that are never sensed, findings that never become action, and improvements that never reach the teams that need them. The engagement builds the missing steps.

Cost

An outage costs more than the downtime. It includes unrecovered revenue, customer churn, SLA credits and engineering time spent on recovery, and it can bring regulatory exposure on top. Estimate yours with the cost of downtime calculator.

Regulation

For financial entities in the EU, the Digital Operational Resilience Act (DORA) requires a post-incident review when a major ICT incident disrupts core activities, and, on request, a record of what changed as a result (Article 13). It also requires a digital operational resilience testing programme (Article 24). It has applied since 17 January 2025.

Scrutiny

A major outage now ends up in front of boards, customers, and the press. An internal postmortem is not enough for any of them.

Start with diagnosis.

Book a call