Resilience engineered, not assumed.

Many organisations have a resilience problem they haven't understood yet. The same incidents keep recurring. We help you understand why, and fix it.

Trusted by Amazon Autodesk Steadybit Indicium AI Uptime Labs

You'll recognise this.

01 The same types of incidents keep happening despite doing incident analysis and service reviews
02 You invested in different practices but can't tell if it's actually making you more resilient
03 The learning never seems to spread across the organisation
04 You're spending more time fighting fires than improving your systems
05 You know something's broken organisationally, but can't pinpoint what
06 Leadership is asking "why does this keep happening?" and you don't have a good answer
07 You're deploying AI into operations but can't tell if it's making things better
08 You're betting on AI to improve operations but you don't know what happens when it's wrong
All these point to the same underlying problem. The organisation has stopped learning.

Start with diagnosis,
then build from there.

First we find out why the same incidents keep happening. Then we fix what's causing them.

How we work, in detail →

A year of this work at Autodesk →

Your organisation keeps having the same incidents despite doing incident reviews, chaos engineering experiments, architecture and service reviews. Something organisational is broken but you can't articulate why.

We embed with your teams to see how the work really happens. We watch how teams interact, what pressures they're under, and where practice and policy don't match.

  • Embedded observation of your practices (incident reviews, GameDays, ORRs, chaos experiments, etc.)
  • Interviews across engineering, operations, and leadership
  • Analysis of your feedback loops and organisational patterns limiting resilience
  • Written report with prioritised, actionable recommendations
  • Readout with your leadership team
2–3 months · available on its own

Full details →

You want to change how your organisation learns in order to improve its resilience capability.

We work with your teams for six to twelve months, often on multiple projects at once depending on how fast we can change your organisation (it takes time). Each work stream is done so your team can take over once we leave.

We first focus on operational excellence and resilience to reduce the firefighting and free capacity for proactive work. When a topic needs a specialist, we bring one in from the Collective.

  • Incident analysis
  • Operational readiness reviews
  • Change management
  • On-call and incident command
  • Observability and tooling
  • Availability measured around critical customer workflows
  • Chaos engineering and GameDays
  • Ownership and roles
  • A shared resilience vocabulary
  • Leadership support, up to the CEO
6–12 months · outcomes-based

Full details →

No prescriptive checklists. We study your organisation's culture first

Founder-led. The person you talk to is the person doing the work

Diagnosis before prescription. We don't sell solutions to problems we haven't seen

What you need to hear, not what you want to hear

Ready to start? Book a call

The work speaks
for itself.

Resilium Labs is led by Adrian Hornsby, a former Principal Engineer at AWS and author of Why We Still Suck at Resilience, now working alongside a curated collective of specialists. Here is what people who have worked with him say.

On credibility
"More often than not, 'consultants' can talk the talk, but cannot walk the walk. If you want to improve the resilience of your systems and operations, Adrian has proven that he can deliver. He is an educator at heart, with in-depth knowledge based on real experience."
WV
Werner Vogels
VP & CTO, Amazon
On the approach
"Adrian doesn't come to you with a prescriptive checklist. Instead, he studies your organisation's culture carefully to understand deep underlying contributing factors that impact resilience. Be prepared for what you need to hear, not what you want to hear. But fear not - Adrian understands human psychology and delivers his insights in a respectful and constructive manner that drives effective and sustainable change. He is an accelerator for organisational learning and improvement."
JN
Jason Niemczyk
Resilience Architect, Autodesk
On domain expertise
"Adrian has now become the go-to independent expert in this space. Most companies don't realize that a good resilience program will speed up their time to market for everything else, and Adrian can help you get there."
AC
Adrian Cockcroft
Tech Advisor, Former VP Architecture, AWS
On socio-technical depth
"Adrian has an exceptional ability to understand the deep interplay between people, teams, and the complex technical problems they are trying to solve. He navigates highly complex socio-technical systems with ease and helps organisations focus on what truly matters. Adrian has a rare talent for enabling teams to work more closely together and build systems that are not only reliable, but resilient by design."
BW
Benjamin Wilms
CEO, Steadybit
On practical impact
"Adrian brought a blend of deep expertise and practical insight to our team. He didn't just teach resilience patterns, he challenged our engineers to think differently about operational excellence and how to design systems with failure in mind. He was engaging, thought-provoking, and left the team with actionable ways to improve how we build and operate software."
SB
Steve Bryen
CTO, Indicium AI
On what outlasts the engagement
"Adrian didn't just advise our team — he helped us build a discipline the organization can now run on its own. A year in, we have a shared vocabulary, a repeatable way to measure what customers actually experience, and a learning loop that turns every incident, GameDay, and review into something the whole company gets smarter from. That's the kind of engagement that keeps paying off long after it ends."
PH
Pat Hughes
Director, Incident Response & SRE, Autodesk
Cover of Why We Still Suck at Resilience by Adrian Hornsby

Why We Still Suck at Resilience

“What accumulates isn’t always what matters.”

Chapter 3 — Why Practices Fail to Build Learning

Your organisation does incident reviews, runs GameDays, and practises chaos engineering. So why do similar incidents keep happening? This book explains why, and what to do about it.

Ready to find out what's actually broken?

The best first step is a conversation to understand your current challenges and resilience goals. We'll help you figure out which step in the journey makes sense for your organisation.

Book a call