The Importance of a Proven Business Continuity Plan
Information technology disaster planning plays a crucial role in making sure a business can still operate after a facilities catastrophe, a natural disaster, or any other serious disruption. Failing to prepare can mean a revenue standstill, profit losses, and major headaches during recovery. As a precaution, most companies of any size draw up a business continuity plan, or BCP, which provides guidelines for responding to a variety of problems. The IT continuity plan is just one piece of that puzzle, and it has to share the same objectives as the company’s comprehensive BCP. I have seen IT plans written in isolation that would restore the servers beautifully while the business itself had no desks, no phones, and no answer for clients. Recovery that is not coordinated with operations is not recovery.
In the event of a disaster, the continued operation of a company depends on its ability to replicate vital IT systems and data. A continuity plan spells out how the company prepares, what the response looks like in the first hours, and what steps restore operations after that. The word “proven” in this post’s title is doing the heavy lifting. A plan that has never been tested is a theory, and in my experience the first test of any plan surfaces problems nobody anticipated on paper: the backup tapes that restore slower than anyone calculated, the vendor contact list with two departed employees on it, the generator that runs but was never load-tested with the actual equipment attached, the single person who knows the storage array password and happens to be on a plane. Periodic testing is imperative, and the tests should be inconvenient enough to be honest. A tabletop walkthrough is a fine start; actually restoring a critical system from backup, on a schedule, is what tells you the truth. If you have never timed a full restore of your most important database, you do not know your recovery time; you know your hope.
The components vary by organization. Facility backup generators, off-site data duplication, failover ISP and telephone connectivity, and even a redundant satellite site can all be part of the picture. In a business built around call center operations, that last pair matters more than most: if the phones and the connectivity are down, the business is down, whatever the servers are doing. Off-site data protection deserves particular scrutiny. Tape rotation to an offsite location is still the workhorse for many of us, and it works, but the plan has to account for the time it takes to recall tapes and perform the restore. Replication to a secondary site shortens that window considerably for the systems that justify the cost. Not every system does, which is the point of the exercise.
That is really the discipline underneath all of this: not every system needs the same protection, and pretending otherwise makes plans expensive and unread. Whether a business runs 24/7 or could function for a day on last week’s data, each plan is unique to the organization. The useful questions are plain ones. How long can each business function be down before the damage is serious? How much data can we afford to lose: an hour, a day, a week? The answers come from operations and finance, not from IT, and they drive everything else: what gets replicated, what gets restored first, and what waits. Writing those priorities down before the disaster is the whole value of the plan, because during the disaster nobody thinks clearly and everybody’s system is the most important one.
One more thing the plan has to cover: people. The plan itself must live somewhere reachable when the building is not, more than one person has to be able to execute every step of it, and someone has to own communicating with employees and clients while the technical recovery is underway. Silence during an outage costs more goodwill than the outage itself.
Have you reviewed or tested your business continuity plan lately? After all, there is no such thing as a benign disaster.