I don't know the details. Nick Sivo is in charge of this stuff, and he'll post something about it. I know he thinks the root of the problem was a disk failure. The server got wedged, and when we rebooted, the file system was corrupted. I'm not sure exactly why it took so long to restore. I was out of town the whole time this was happening.
The reason we lost so much data was that we only do nightly backups. That seemed enough when we started. Now that HN is a bigger part of more people's lives, we'll make more of an effort to make it proof against this sort of problem.
Another point to make is that, even with a lot of data loss, it at least would be good in the future to skip a bunch of post identifiers (potentially even just saying "well, we certainly didn't use 100,000 of them, so I skipped up to 7100000") in order to try to not cause data that used to be there, and may have been archived by various sources including Google Cache or in the various reader clients people use on various devices , to suddenly have been swapped out by different posts leading to wide-spread cache corruption (and further confusion or loss). I mean, even just for the sake of people who may have posted links to this content in various places: maybe they wrote an article or posted a tweet/comment somewhere referencing how great/horrible a post was, and now suddenly it is saying something entirely different and potentially quite awkward ;P.
As a quick example, for those still not sure what I mean: what used to be an article about Python 2/3...
I'll post something more detailed tomorrow, but in terms of data loss, we went down at 2014-01-05 16:10:29 PST and restored a backup from 2014-01-05 01:00:00 PST.
It seems to be valued much more by the community than those who manage it - this is a problem IMO as there's then an inbalance in the desire of the community to have an optimal tool and the desire of the management which appears to be more to provide something akin to an MVP.
Of course the community is diverse and my view is not necessarily representative of any type of majority.
Thanks for the update. I don't think all that much was lost, and a day to restore from a disk failure ain't bad. Please thank Nick and whoever else worked overtime to get things back up and running!
The reason we lost so much data was that we only do nightly backups. That seemed enough when we started. Now that HN is a bigger part of more people's lives, we'll make more of an effort to make it proof against this sort of problem.