Change Management Windows — The Star Trek Way

More than a decade ago, I wrote a blog post about Change Management windows using an example from Star Trek: The Next Generation.

A lot has changed in IT since then. Our technology is different. Our deployment methods are different. Agile, DevOps, automation and cloud have changed how quickly we can introduce change.

But when I went back and looked at that old post, I realized something.

A lot of the basic Change Management lessons still apply.

And apparently, the idea I was talking about even has a name: the Scotty Principle.

The idea comes from the running joke that Scotty would tell Captain Kirk a repair would take considerably longer than he actually expected it to take. When he completed it early, he looked like a miracle worker.

There is a great version of this in Star Trek: The Next Generation. when Scotty meets Geordi La Forge.

Geordi tells Captain Picard that a repair will take about an hour. Scotty asks him how long it will really take.

Geordi tells him: an hour.

Scotty can't believe it. How is Geordi ever going to maintain his reputation as a miracle worker if he tells the captain how long the work will actually take?

 

It's funny because there is probably a little bit of Scotty in most IT organizations.

Someone asks how long something will take. We think it will take two hours. Then we start wondering:

  • What if something goes wrong?
  • What if testing takes longer?
  • What if we need to back it out?

Maybe we should ask for four hours.

There is nothing wrong with contingency. In fact, there should be some.

But there is a difference between planning for risk and simply padding an estimate.

That difference is where Change Management comes in.

How long should a change window actually be? This was the question behind the original post, and it is still a good question.

  • If you ask for too little time, you risk running beyond the approved window, rushing validation or handing a service back late.
  • If you ask for too much, the business may push back on the outage. Do it often enough and people may also stop trusting your estimates.

Neither is particularly good.

The answer isn't to find the perfect amount of padding. The answer is to get better at understanding the change. That means asking some fairly basic questions:

  • How long should the implementation actually take?
  • How much time do we need to test and validate it?
  • What could reasonably go wrong?
  • At what point do we stop trying to fix the implementation and decide to back out?
  • How long will the backout take?
  • And when does the business actually need the service available again?

That last question is particularly important.

A technically convenient change window isn't necessarily a good business change window.

 

Change Management is really about risk

This is one of the things we spend time on in integratedITSM™ Essentials.

Change Management isn't about completing a form, getting an approval or putting every change in front of CAB. Those are activities that might support the process. The real objective is to manage the risk associated with introducing change while still allowing the organization to make the changes it needs.

That's an important distinction.

A mature Change Management process shouldn't make change difficult. It should help us understand how much control is appropriate for the level of risk we're taking.

A routine, repeatable change that we've successfully performed 50 times shouldn't necessarily require the same scrutiny as a complex change to a critical production service.

And we should have the data to know the difference.

 

Every change should teach you something

This is probably the part I would add most strongly to what I wrote more than a decade ago.

Your change records shouldn't just document what you plan to do.

They should help you learn.

Suppose we planned:

  • Implementation: 60 minutes
  • Testing and validation: 30 minutes
  • Contingency: 30 minutes
  • Backout: 60 minutes

Then we perform the change.

  • Implementation actually takes 35 minutes.
  • Testing takes 50.
  • Everything works and we return the service within the agreed window.

Great.

But what happens next? Too often, we close the change and move on.

We shouldn't.

We should capture what actually happened and understand:

  • Why did implementation take less time?
  • Why did testing take longer?
  • Was the backout estimate realistic?
  • Did we discover something that would make the next implementation easier?
  • Was there a step missing from the implementation plan?

That information becomes part of the organization's knowledge.

After performing the same type of change 10, 20 or 50 times, our estimates should be getting better.

If they aren't, we aren't learning from the process.

 

This is also where metrics become useful

We tend to measure things like change success rate. That's useful, but it doesn't tell the whole story. You could have a 98% change success rate and still have a pretty poor Change Management process. we really need to understand:

  • How many changes exceeded their approved windows?
  • How accurate were the original estimates?
  • How many resulted in incidents?
  • How many needed to be backed out?
  • How many required emergency intervention?
  • Are particular types of changes consistently taking longer than expected?
  • Are certain teams, applications or services generating more failed changes?
  • And are we getting better? - That's the question I really care about.

A metric shouldn't just tell me what happened.

It should help me decide what to do differently next time.

That's why measurement and continual improvement are important parts of integratedITSM. We're not collecting numbers because somebody wants a dashboard. We're collecting information that helps us improve how the system works.

 

Sometimes the answer is a smaller change

This was another point from the original article that still holds up. If you can't find a reasonable window for a large change, maybe the answer isn't a larger window. Maybe it's a smaller change.

  • Can we break the implementation into components?
  • Can we change one thing, validate it and then move to the next?
  • Can we reduce the number of variables we're changing at the same time?
  • Can we automate part of the implementation?
  • Can we reduce the blast radius?

There is another benefit here that sometimes gets overlooked - Troubleshooting.

If you change five things at once and something breaks, you now have five potential causes.

If you change one thing, validate it and then move on, you have a much better idea where to start looking when something doesn't behave as expected.

That's not bureaucracy.

That's risk management.

 

Change Management doesn't work by itself

This is another area where my thinking has evolved since I wrote the original article. We sometimes talk about ITSM processes as though they live independently.

They don't.

A change might exist because Problem Management identified the underlying cause of recurring incidents.

Incident Management may tell us what happened the last time a similar change was implemented.

Configuration Management can help us understand what components and services could be affected.

Release and Deployment Management helps us introduce the change into the environment.

Service Level Management helps us understand the commitments we're trying to protect.

And Business Relationship Management helps us understand what the business can actually tolerate.

 

This is one of the central ideas behind integratedITSM.

The individual processes matter.

But understanding how they work together is where service management becomes much more useful.

 

So, what happened to the Scotty Principle?

I still like it.

And there is actually some wisdom buried inside it.

Scotty understood uncertainty. He knew things didn't always go according to plan, and he allowed himself room to deal with that.

That's not a bad instinct.

The problem comes when contingency becomes arbitrary padding.

Good Change Management should replace that guesswork with knowledge.

  • Use previous changes to improve your estimates.
  • Understand your risk.
  • Know your backout time.
  • Build in appropriate contingency.

Measure what actually happened. Learn from it. Then use that information the next time. Do that consistently and something interesting happens.

You don't need to look like a miracle worker.

You become predictable.

And when it comes to Change Management, being predictable is probably more valuable to the business than being a miracle worker anyway.

That's a big part of what we teach in integratedITSM™ Essentials: not simply what the individual ITSM processes are, but how they work together, how we manage risk, how we measure performance, and how we use what happens today to make the process better tomorrow.