Every single safeguard was there. It was just that someone had economized on each one. And economizing adds up, until one night presents the whole bill.
It was about a live financial system, a continent away, in Mexico. It landed on my desk because the local team wanted to go after the operator. The operator was threatening to stop running backups until its contract was finally renewed. That contract had been expired for twelve months, and no one in the group had responded to the renewal requests. A backup is highly critical, so I was asked to take a look.
What followed was a textbook chain. Why is the backup not running? Ask the software maker. Answer: no valid license anymore, not renewed. So bought ad hoc. The software maker says: database problem. Support from the database maker? No longer provided, no contract. Bought ad hoc. The database reports: blocks can no longer be written, the disk is defective. The hardware maker? No maintenance contract. Bought ad hoc. And indeed, we need new disks. Except that, to migrate to a new system, we need the contents of exactly those defective blocks.
Block by block
In the end, for a lot of money, we brought all three makers to one table and, record by record, rescued the still-readable blocks into a new database with extraction tools. A remainder was lost. Then came the walk to the Mexican tax authority: records were missing, we could not restore them. The authority was open about it, asked for a documented estimate, and that settled the matter.
The moment it became clear to me how deep the hole was, I smiled. After enough SAP years, escalating clients and sleepless nights, something like this no longer shocks you.
We averted it with the local team and us from Europe. The real reason for the situation lay deeper, in an unhealthy structure, and only that made the disaster possible. The two sites were far apart, there was no central steering of the operator, locally people preferred to argue, and the European side of the operator had little real leverage, no matter how loudly we escalated. What was missing is the oldest rule there is: save in good times and you will have it in need. Whoever builds load-bearing structures in normal operation has something to stand on in the crisis.
Not long after, the responsible CIO left. The connection was clear. In his own interest in transparency, he should have seen the lapsed contracts and the operator structure and fixed them with procurement. It went well once more. That was pure luck, not merit.
So never check your fallback in isolation, but as a chain: backup, license, maintenance, spare part. One economized link is enough to make all the others worthless.
An economized maintenance contract is invisible. The bill for all the economized ones together comes on a single night, the one on which everything else has already gone wrong.