Incident Response

    IT Recovery Is Not Business Recovery: What the Restore Button Does Not Fix

    Jeff SowellSeptember 10, 2026
    IT Recovery Is Not Business Recovery: What the Restore Button Does Not Fix

    InformationWeek recently ran a piece asking a deceptively simple question: when is a business really recovered from a cyberattack? I was one of the people they interviewed, and my answer made it in: safety and revenue first. Everything else is secondary.

    A quote is a sentence. The thinking behind it takes more room than a trade article allows, and the piece deliberately left a few hard questions on the table. This is the longer answer, drawn from running recovery planning and incident exercises for companies whose boards ask exactly these questions.

    The core distinction the article got right

    Restoring systems is a technical milestone. Recovery is a business outcome. Those are different finish lines, and most organizations only plan for the first one.

    Here is the test I use: if your incident commander announces "all systems restored" and your CEO asks "so are we back?", does anyone in the room have a defensible answer? At most companies the honest response is silence, because nobody defined what "back" means in business terms. Orders shipping at normal volume. Payroll running. Customers who did not quietly start a vendor search. None of that appears on an infrastructure dashboard.

    Question one: speed versus hardening

    The article asked how organizations balance recovery speed against security hardening, and it is the right tension to name. After an intrusion, the fastest path back is restoring what you had. The problem is that what you had is what got breached, and in ransomware cases the attacker's access often predates your oldest clean backup.

    The answer is not a universal policy. It is a decision made per system, before the incident, on two axes: how much revenue or safety depends on this system per hour, and how likely is it to be a re-entry path. A customer-facing booking system might come back fast in a degraded, closely watched state because every hour costs real money. The identity provider that was the attacker's front door does not come back until it is rebuilt, because restoring it fast means running the incident twice.

    Writing those calls down ahead of time is the whole game. In the middle of an incident, every voice in the room argues for speed, and the person arguing for a clean rebuild sounds like the enemy of the business. A pre-agreed decision rule, signed by the executive team on a calm day, is what lets the security lead say no without it becoming a career moment.

    Question two: the gap between IT recovery and business recovery

    The article asked what timeline separates the two. In my experience the honest answer is that IT recovery is measured in days and business recovery in weeks to months, and the ratio surprises every executive team the first time they see it.

    Systems come back and then the queue starts: the order backlog that built up during downtime, the invoices that did not go out, the reconciliation work when half your records were restored from Tuesday's backup and the other half from Thursday's. Then the slower layers: the insurance claim, the customer notifications, the auditors who now want evidence about everything, the sales cycle that stretched because a prospect read about your incident. Revenue impact routinely outlives the outage by a quarter or more.

    This is why I push clients to define recovery targets for business processes, not servers. A recovery time objective on a database tells you when IT is done. A recovery objective on "we can take and fulfill a customer order" tells you when the business is breathing again. The second number is the one the board actually cares about, and it is the one almost nobody has written down.

    Question three: partial recovery is the normal case

    The article's third open question was how to handle partial recovery, and the framing matters: partial recovery is not an edge case, it is the default. You will run degraded for days or weeks. The question is whether you chose what runs first or discovered it by accident.

    The tool for this is tiering. Every system gets a tier with an explicit recovery order, and the tiers are set by business consequence, not by which team is loudest. In our client work, the top tier is reserved for anything touching life safety or direct revenue capture, with recovery targets measured in minutes, not hours. The bottom tier is everything the company can live without for a month, and being honest about how much lands there is what makes the top tier achievable.

    Degraded-mode playbooks belong next to the tiers: how do we take orders when the ERP is down, how does the front desk work when the access system is offline. Paper works. Spreadsheets work. Not knowing is the only thing that does not work.

    The exercise that answers all three questions

    In the article I recommended exercises where you deliberately take systems away, because the workarounds and vendor dependencies you discover are the ones that would kill you in a real crisis. Let me describe what that looks like, because it is different from the tabletop most companies run.

    A typical tabletop is a conversation: here is a scenario, what would you do. Useful, but it tests memory of the plan, not the plan itself. The removal exercise is different. Pick a critical system. Announce that as of now it does not exist, and neither does its vendor's support line. Then watch. Who reaches for a workaround that depends on the same identity provider that is also down? Which "fully redundant" process turns out to route through one person's laptop? Which vendor dependency did nobody know existed until the room needed it?

    Every one of those discoveries costs almost nothing in an exercise and a fortune in a real incident. And the findings feed straight back into the three questions above: the speed-versus-hardening calls get realistic, the business recovery timeline gets honest, and the tier order gets corrected by evidence instead of opinion.

    What to report upward while it burns

    One more thing that did not fit in a quote. During recovery, executive communication should carry three numbers: what percent of normal business capacity we are operating at, what each day at that capacity costs, and the realistic date for the next capacity step. Database restoration status is not an executive metric. Capacity, cost, and date are, and teams that report that way keep their leadership's trust through the worst weeks of the incident.

    Frequently Asked Questions

    What is the difference between IT recovery and business recovery?

    IT recovery means systems are restored and running. Business recovery means critical operations are producing at normal capacity: orders flowing, payroll running, customers retained. IT recovery is typically measured in days; business recovery in weeks to months, because backlogs, reconciliation, insurance, and customer trust all lag the technical restore.

    Should we restore fast or rebuild clean after a ransomware attack?

    Per system, decided in advance. Systems with high hourly revenue or safety impact and low re-entry risk come back fast in a monitored, degraded state. Systems that were the attack path get rebuilt clean, because restoring them fast risks running the incident twice. The decision rule should be agreed and signed off before any incident.

    What is a system-removal exercise?

    An incident exercise where a critical system is declared gone, including its vendor support, and the team must actually operate without it rather than talk through a scenario. It surfaces hidden workarounds, single points of failure, and unknown vendor dependencies at exercise cost instead of incident cost.

    What should executives be told during cyberattack recovery?

    Three numbers: current percent of normal business capacity, cost per day at that capacity, and the realistic date of the next capacity improvement. Technical restoration milestones belong in the engineering channel, not the board update.

    How do we start if none of this is written down?

    Tier your systems by business consequence, put a recovery target on your top revenue-producing process rather than on servers, and schedule one removal exercise. Those three steps take weeks, not quarters, and they convert your incident response plan from a document into a capability. A vCISO engagement is one way to get them done with experienced hands.

    Related reading: our guide to incident response planning, and if you want a fast read on where your program stands today, the free cybersecurity assessment takes a few minutes. The original InformationWeek piece is here.

    incident responsebusiness recoverydisaster recoverycyberattacktabletop exercise

    Related from the BlueRadius Library

    Sourced posts on adjacent topics, ranked by tag overlap.

    Have a security story worth telling? We publish practitioner guest articles.

    Write for us

    Take the Next Step

    Ready to Strengthen Your Security Posture?

    BlueRadius delivers Fortune 500-grade protection for mid-market companies — virtual CISO leadership, 24/7 managed security, and compliance programs that actually close deals. Let's talk.