How we work.
We find out what is going on before we change anything.
We embed with your teams. We sit in on incident reviews, GameDays, chaos experiments, and operational readiness reviews, and we interview people at every level, from the engineers on call to the CEO.
The diagnosis names the patterns behind your recurring incidents. Most of them sit outside the engineering practices: how leaders behave under pressure, what gets rewarded, and how software and data are built and operated.
The diagnosis is where the work starts. We work with your teams for six to twelve months on the conditions it found. When an incident happens during the engagement, we analyse it together as it happens. Every practice we build is designed so your team runs it after we leave.
We ask what conditions made the failure possible, and why the decisions involved made sense to the people who made them at the time. We look for what surprised people, and for the weaknesses that haven't caused an incident yet. The analysis holds multiple teams accountable without blaming any of them, and puts the gaps in the system rather than in individuals.
No prescriptive checklists. We study your organisation's culture first
Founder-led. The person you talk to is the person doing the work
Diagnosis before prescription. We don't sell solutions to problems we haven't seen
What you need to hear, not what you want to hear
Every organisation has a gap between what it thinks happens in its systems and what actually happens. The gap reopens with every change. Learning is how the organisation closes it, and it happens in three steps.
Sense the gap: incidents do it on production's schedule; load testing, chaos engineering, and GameDays do it on yours. Understand it: incident analysis after the fact, operational readiness reviews before the next change. Improve: change the system, the practices, and what people believe. Then sense again. Resilience is closing the gap at least as fast as it reopens, and before production does it for you.
Every practice can run in two modes: as a ritual that performs diligence, or as a mechanism that feeds what it finds back into the loop. Which mode you get depends on what sits under every practice: psychological safety, the right incentives, and leadership support. That is the bedrock. Without it, no practice works.
The assessment finds where your loop is broken: gaps that are never sensed, findings that never become action, and improvements that never reach the teams that need them. The engagement builds the missing steps.
Cost
An outage costs more than the downtime. It includes unrecovered revenue, customer churn, SLA credits and engineering time spent on recovery, and it can bring regulatory exposure on top. Estimate yours with the cost of downtime calculator.
Regulation
For financial entities in the EU, the Digital Operational Resilience Act (DORA) requires a post-incident review when a major ICT incident disrupts core activities, and, on request, a record of what changed as a result (Article 13). It also requires a digital operational resilience testing programme (Article 24). It has applied since 17 January 2025.
Scrutiny
A major outage now ends up in front of boards, customers, and the press. An internal postmortem is not enough for any of them.