Reliability as a human system
Reliability depends on clear ownership, practiced recovery, and people who raise concerns early.
When I review reliability, I start with technical behavior such as availability, recovery time, and capacity. Then I look behind those measures. I want to know who decides what matters, how work changes hands, when people escalate, and whether they can describe risk plainly.
That second view is why I treat reliability as a human and organizational problem. A component may fail for a technical reason, but choices made much earlier shape the effect. An unclear owner can delay the response. A routine exception may have become invisible through repetition. Someone may have noticed a concern and decided that raising it was too costly.
Build structure before asking for more energy
A team may appear fast by relying on overtime, concentrated knowledge, and repeated escalation to its most experienced people. That approach can clear an immediate obstacle. It also hides queues and dependencies that return with the next urgent request.
I have found that a structured team moves faster over time. When ownership and decision rights are clear, routine choices can proceed without another meeting or a search for the right person. A useful runbook gives someone a place to begin under pressure. Smaller, reversible changes make a wrong answer less expensive, and shared standards keep familiar work from requiring a new negotiation.
The purpose of that structure is to protect attention for situations that are genuinely new. Explicit expectations help people recognize an exception sooner and spend less time reconstructing basic context.
I treat any process that depends on an unwritten rule as a reliability risk. Memory varies, people change roles, and instructions are sometimes misunderstood. A sound operating model accounts for those ordinary conditions instead of relying on individual stamina to overcome them.
Make it safe to raise a concern
People need ways to say, “I do not understand this” or “The plan no longer matches reality.” Psychological safety can lower the personal cost of speaking, while candor helps the team examine the concern without avoiding discomfort.
My management standard is to separate a person’s worth from the quality of a decision or result. I want people to challenge weak reasoning and honor clear commitments. I also want concerns raised while options remain, rather than after a quiet doubt has become part of an incident.
A manager’s response to bad news influences the next report. I pay close attention to my first reaction. Punishing the person who surfaced a problem teaches everyone else to wait. Excusing the miss would make the standard harder to trust. I try to understand the sequence before assigning blame, contain the consequence, and then decide where accountability or the surrounding conditions need to change.
I use the same candor when talking about capacity. When demand exceeds what a team can safely carry, naming the limit early creates choices. Work can be narrowed, sequenced, or deferred. Concealing the limit behind extra effort postpones the decision while increasing operational risk.
Practice before the incident
An incident reveals much of the operating system already in place. There is little time during a disruption to build trust, establish decision rights, or invent a shared vocabulary. People draw on what they have practiced and what they can find.
Preparation can remain plain. I begin with which services matter most. Everyone involved should be able to find the escalation path and know who has authority to direct the response. We rehearse recovery often enough to expose stale assumptions. I also ask whether monitoring points toward a decision or merely adds noise. Afterward, the review should reconstruct the sequence and contributing conditions without collapsing the explanation into one person’s mistake.
Because a procedure cannot anticipate every failure, I use it to help the team orient. What is affected? Who is deciding? What must be protected first? Which evidence will show that recovery is holding? Those questions create a stable frame for judgment when the details are unfamiliar.
I apply the same preparation to automation. Before removing a manual handoff, I want to understand its purpose, likely failure modes, stop conditions, and validation step. With that clarity, automation can reduce repetitive work and improve consistency. Without it, the automated process may simply produce errors more quickly and at a larger scale.
What I look for between incidents
Complex systems will fail. My job is to keep risk discussable before a disruption and respond without drama or evasion when one occurs. What we learn afterward should change the next response.
Uptime alone does not tell me whether an operation is healthy. I also look for recurring failure patterns and instructions that no longer match the work. Knowledge held by too few people concerns me, especially when an exhausting on-call practice or repeated heroics are covering the gap. I treat these as warning signs because each can reduce the organization’s ability to adapt.
In reliable teams I have worked with, people understood their roles and could tell the truth about what they saw. They moved quickly when speed mattered because they had practiced, not because sustained exhaustion had become normal.
That is what I want to leave behind: a team that can recognize a problem, decide what to do, and recover without waiting for the same expert to rescue it. If the operation becomes easier to understand and less dependent on heroics, I consider that real reliability work.