Most managers plan their team’s week as if the whole week were available. Five people, forty hours each, two hundred hours to spend. You assign the roadmap work against that number, feel good about the plan on Monday morning, and then watch it fall apart by Wednesday. The plan didn’t fail because your estimates were bad. It failed because a large slice of those two hundred hours was never yours to assign. It belonged to work nobody scheduled.
A New Relic study of 1,700 IT professionals put a number on that slice. Teams spend roughly a third of their time responding to disruptions rather than doing planned work, and the median organization absorbed 280 hours of downtime a year at a cost that ran into the hundreds of millions across the sample (New Relic, via DevOps.com, 2024). A third. If you planned your team’s week against 100 percent of their hours, you overcommitted them by roughly 33 percent before anyone touched a keyboard.
I have watched this pattern play out for more than twenty years in IT operations and in the fractional COO engagements I run now. The teams that miss their commitments quarter after quarter are rarely lazy or slow. They are buried in work that shows up unannounced, and their manager is still planning as if that work doesn’t exist.
The four buckets your team’s work actually falls into
The clearest way I’ve found to talk about this comes from Gene Kim’s The Phoenix Project, which sorts all work into four types: business projects, internal projects, changes, and unplanned work (IT Revolution summary). The first three are things you decide to do. You can see them, plan them, sequence them. Unplanned work is the outlier. Kim calls it recovery work, and his definition is worth sitting with: it is urgent, reactive work that pulls you away from your goals, and it is almost always the symptom of something else that broke.
That last part is the piece managers miss. Unplanned work is not weather. It is not a random act of nature you simply endure. It is the bill coming due for a bug you shipped, a process that was never documented, a system only one person understands, a handoff that dropped. When your team spends Thursday firefighting, Thursday’s fire was lit by a decision made weeks earlier.
Distinguish this from two things I’ve written about before. This is not the same as toil, the repetitive manual work that scales with your service and grinds people down. Toil is predictable and boring. Unplanned work is unpredictable and loud. And it’s different from the utilization trap, where you book people to 100 percent and lose all your flow to queuing. You can be under-utilized on paper and still lose your week to fires. These problems compound each other, but they are not the same problem.
Why it costs more than the hours it burns
If unplanned work only cost you the hours it directly consumed, you could just pad the plan and move on. The reason it’s so corrosive is that it damages the planned work it interrupts, and that damage doesn’t show up on any timesheet.
Gloria Mark’s research at UC Irvine found that after a single interruption, a worker takes an average of 23 minutes and 15 seconds to return to the original task at the same depth of focus (Fast Company on Mark’s work). She named the mechanism attention residue: part of your mind stays stuck on the interrupting task even after you’ve handed it off. Three interruptions and you’ve lost the better part of an hour that never appears in any tracking tool, because on paper that hour was “spent on the project.”
So the real math is worse than the survey number. A team losing a third of its hours to unplanned work is also losing a chunk of the remaining two-thirds to the recovery tax on either side of every interruption. This is why teams in heavy firefighting mode feel like they’re running flat out and shipping nothing. They are. The reactive work eats the hours, and the residue eats the focus that was supposed to make the surviving hours productive. I’ve covered the mechanics of this in more detail in the piece on the switching tax, but the operational point here is simpler: an interruption is never just the length of the interruption.
Measure the ratio before you try to manage it
Here is the uncomfortable first step, and it’s the one most managers skip. You almost certainly do not know what percentage of your team’s time goes to unplanned work. You have a feeling. The feeling is usually wrong, and it’s usually low.
For two weeks, tag the work. It doesn’t need a fancy system. A shared column in whatever you already use, or a simple tally, with one question against each task: was this on the plan Monday morning, or did it arrive after? Two buckets, planned and unplanned. Count the hours in each at the end of the two weeks.
I ran this exercise with a client team last year that swore they were “mostly on the roadmap.” When we actually tagged it, unplanned work came in at 41 percent. The manager had been planning sprints against full capacity and then treating the perpetual overrun as a team performance problem. It was never a performance problem. It was a measurement problem. Once the number was visible, the conversation changed from “why are you always behind” to “we have a team running on 59 percent of its capacity, and we’ve been pretending it’s 100.”
The number itself is diagnostic. Under 15 percent and unplanned work is normal operational noise; don’t over-engineer it. Between 15 and 30 percent, plan for it explicitly. Above 30 percent, you have a system that is generating fires faster than your team can put them out, and no amount of better prioritization will fix that. You have to go after the sources.
Put a ceiling on it, then defend the ceiling
Google’s Site Reliability Engineering practice handles this with a rule I’ve borrowed for plenty of non-engineering teams: operational and reactive work is capped at 50 percent of a person’s time, and when it runs over, the excess gets pushed back to the team that generates it until the load drops (Google SRE, “Eliminating Toil”). The cap matters less than the mechanism. Without a ceiling, reactive work expands to fill every available hour, because it’s always urgent and urgent always wins against important.
For most management teams, a hard 50 percent cap is aspirational. But the principle holds at any threshold. Decide what fraction of your team’s capacity you are willing to leave open for unplanned work, plan your committed work against what’s left, and treat a breach of that line as a signal to act rather than a reason to ask people to work later. If you’ve measured 30 percent unplanned, you plan roadmap work against 70 percent of capacity. Not 100. Not “70 percent plus a stretch goal.” Seventy.
This is the part that requires spine, because it means committing to less visible output in exchange for actually hitting what you commit to. Every manager I know finds this hard the first time. The ones who do it stop having the same demoralizing conversation every retro about why the plan slipped again. The plan stops slipping because the plan finally matches the capacity.
The fire tells you where to look
Capping unplanned work buys you room. It doesn’t reduce the fires. To reduce them you have to treat each one as evidence, because Kim was right that unplanned work is a symptom.
Keep a lightweight log of what actually interrupted the team, not to assign blame but to find the repeat offenders. After a month you will see the pattern almost every operations team has: a small number of sources generate most of the reactive load. One brittle integration. One process with no documentation, so every edge case becomes an escalation. One system only a single person understands, which is its own key-person risk waiting to become a full crisis. The fires cluster.
That clustering is good news, because it means you don’t have to fix everything. You have to fix the two or three sources generating 80 percent of the interruptions. That’s where root cause analysis earns its keep, and it’s why the log matters more than the firefighting heroics. A manager who celebrates the person who stayed until 9pm to put out the fire, but never asks why the fire started, is training the team to keep the building flammable.
The same discipline applies at the front door. A surprising share of what looks like unplanned work is really unmanaged intake: requests that skipped the queue, favors that became commitments, “quick questions” that ate an afternoon. Tightening how work enters the team, which I’ve written about under work intake, converts a chunk of the reactive pile back into planned work you can actually sequence.
The shift that changes everything else
The managers who get out of permanent firefighting mode all make the same mental move. They stop treating unplanned work as the exception that ruins the plan and start treating it as a permanent line item in the plan itself. It has a size. That size is measurable. It has sources. Those sources are fixable. And the capacity it consumes is not available for anything else, no matter how badly you want it to be.
Your team’s real capacity was never the number of people times the hours in a week. It’s whatever survives the firefight. The only question is whether you’re planning against that real number, or against the fiction on the org chart. Measure it this week. The number will be higher than you think, and knowing it is the whole game.