← Work
Case study · Autodesk

Building a resilience practice at Autodesk

Client
Autodesk
Industry
Design and make software, cloud platforms
Engagement
12 months, embedded advisory
Focus
Operational readiness, availability measurement, chaos engineering, incident analysis, organisational learning

“Adrian didn’t just advise our team — he helped us build a discipline the organization can now run on its own. A year in, we have a shared vocabulary, a repeatable way to measure what customers actually experience, and a learning loop that turns every incident, GameDay, and review into something the whole company gets smarter from. That’s the kind of engagement that keeps paying off long after it ends.”

— Pat Hughes, Director, Incident Response & SRE, Autodesk

About Autodesk

The world’s designers, engineers, builders, and creators trust Autodesk to help them design and make anything. From the buildings we live and work in, to the cars we drive and the bridges we drive over. From the products we use and rely on, to the movies and games that inspire us. Autodesk’s Design and Make Platform unlocks the power of data to accelerate insights and automate processes, empowering our customers with the technology to create the world around us and deliver better outcomes for their business and the planet. For more information, visit autodesk.com or follow @autodesk. #MakeAnything

The work

Autodesk’s cloud portfolio has grown fast, and with that growth the bar for reliability keeps rising: when a shared platform serves every product, an hour of downtime lands on job sites, factory floors, and production deadlines far beyond the software itself.

Inside Autodesk’s Trust organization, the Site Reliability Engineering (SRE) team had taken on an ambitious mandate to become the company’s resilience experts and enablers, setting the standards and practices that hundreds of service teams build against. The baseline assessment in the first month found a starting point most organisations do not have: active GameDays, blameless debriefs, chaos engineering, and service reviews, already in place. The team had strong engineers, leadership support, and a general sense of direction. What they wanted was calibration in order to produce the maximum learning. When your goal is to deliver world-class availability and reliability, the hardest question about your practices is what good actually looks like, and the fastest way to answer it is to work alongside someone who has seen it built before.

That is what the engagement was for.

Resilium Labs embedded with the team for a year as an advisor. The work was deliberately light on tooling and heavy on practice, vocabulary, and judgment. Four threads ran through it.

A common vocabulary for tradeoffs. Resilience work stalls when the same words mean different things in different rooms. Terms like resilience, reliability, availability, static stability, adaptive capacity, and critical workflow, etc., each had as many definitions as teams using them. We built a resilience lexicon for the organisation, with each concept defined, sourced, and tied to Autodesk’s own context. The point of a lexicon is to navigate the tradeoff the organisation faces, and you cannot do that until both sides of the tradeoff have names. After a few weeks, team members were spontaneously using the engagement’s distinctions in their own discussions, controls versus guardrails, work-as-imagined versus work-as-done, etc.

“Adrian helps me define a meaningful vocabulary that I can share and align on with various stakeholders across the company to achieve clearly defined objectives more effectively.” — Jason Niemczyk, Resilience Architect, Autodesk

Measuring what customers experience. Availability numbers only matter if they reflect what users actually do. The team defined critical workflows for the organisation’s most critical products, a handful of end-to-end journeys that represent core customer value, and rebuilt availability measurement around them: synthetic probes per workflow, conservative aggregation, and a consistent method across organisations so that two teams reporting availability mean the same thing. This work also fed a cost-of-outage model the team can use to size reliability investment against what incidents actually cost, across Autodesk and its customers.

Operational readiness and chaos engineering. The team designed an Operational Readiness Review process for services heading to production, with Resilium Labs advising on structure, reviewer guidance, and the boundary between readiness reviews and the other review mechanisms already in place. Reviews were designed as guardrails that help teams move safely rather than gates that stop them, and reviewers were trained as learning facilitators rather than quality inspectors. On chaos engineering and GameDays, Autodesk already had a practice running. The work there was the difference between validation and discovery. Validation confirms what the team already knows; discovery reveals what it doesn’t, and discovery is where the learning is. The team built experiments around hypotheses worth testing, treated surprises as the valuable outcome, and carried results back into the system instead of into a report.

Building the capacity to learn. Incident analysis, GameDays, chaos experiments, and readiness reviews form one loop: each practice reveals gaps the others test, and what one finds feeds the next. An organisation improves its resilience at the rate that loop turns. Every practice in it can run in two modes: as a ritual that performs diligence, or as a mechanism that feeds what it finds back into the living system.

For incident analysis, Resilium Labs worked with the team on the question sets that reach past the first loop of learning, what failed and how to fix it, into the second, more important one: what made this the reasonable way to work, and what does the organisation change about that. In practice that means “what was the root cause” becomes “what conditions made this possible” and the most useful question “what surprised you.” The later analyses hold multiple teams accountable without blaming any of them, name what worked alongside what failed, and locate gaps in the system rather than in individuals.

The operational readiness reviews got the same treatment, designed so the team being reviewed leaves with learning instead of a pass/fail stamp. The same thinking shaped the weekly operational review the team has just launched, designed so that service teams present their operations in their own words, the week’s important postmortems circulate across the organisation, and the meeting stays a learning conversation, kept deliberately separate from compliance reporting.

“You give us language, ideas, thoughts, and even frameworks we couldn’t easily get otherwise. … I never feel like I’m being sold to or forced to see things your way; it’s always presented in an approachable way that encourages further healthy discussion and ultimately facilitates learning.” — Alan Williams, Senior Principal Engineer, Autodesk

Midway through the engagement, a major cloud provider outage put the work to a real test. The team took ownership of the meta incident analysis, the cross-service review that spans many teams and usually satisfies none of them, and wrote their strongest analysis of the year: multiple teams held accountable, none of them blamed, the gaps located in the system.

What stands at the end

No consultant should be the interesting part of the story a year in. The measure of the engagement is what the team owns as it ends. Some of it is running, some is mid-rollout, and all of it is theirs:

  • A shared resilience vocabulary in active use in design discussions and documents
  • A new framework for incident analyses that surface system-level patterns and circulate as learning artifacts
  • An Operational Readiness Review process designed by the team and in active rollout, with the first pilot review run
  • A clear definition and framework for identifying critical workflows, being applied across products
  • A weekly operational review launched, with the team learning the format as it runs
  • Chaos experiments and GameDays designed around discovery
  • A company-wide reliability and resilience policy drafted by the team, working its way through review and already driving the right conversations

The team set out to become the organisation’s resilience experts and enablers. That is now the job they are doing.

“Over the past year, Adrian has been a trusted advisor to our SRE team. He helped us strengthen our reliability and resilience standards, refine critical workflows around customer value, and shape operational reviews as learning mechanisms. His perspective challenged our assumptions and helped us clarify key trade-offs and turn our ideas into actionable practices.”

— Luca Mitic, Senior Manager, Site Reliability Engineering, Autodesk

About Resilium Labs

Resilium Labs helps engineering organisations build resilience practices that outlast the engagement: operational readiness, availability measurement, chaos engineering, and the learning culture that ties them together. It is founder-led by Adrian Hornsby, former Principal Engineer at Amazon Web Services and author of Why We Still Suck at Resilience.