Skip to content

The news behind the software that runs companies, explained for everyone.

A GitLab engineer deleted about 300 GB by mistake. His own snapshot saved it.

ERP LEADERS desk ·
Video: GitLab

Late on 31 January 2017 he was repairing the copy of GitLab.com's main database. He had said earlier he would log off. The live server was called db1, the copy db2. He meant to wipe db2 and ran the command on db1.

He noticed within a second or two and stopped it. GitLab's incident notes: "Of around 300 GB only about 4.5 GB is left."

The team went for the backups. GitLab later wrote that "out of five backup/replication techniques deployed none are working reliably or set up in the first place."

What worked was a snapshot of production the same engineer had taken by hand about six hours earlier and loaded into staging. Copying it back over slow disks took about 18 hours, streamed live on YouTube. Six hours of database changes were gone: GitLab estimates roughly 5,000 projects, 5,000 comments and 700 new user accounts were affected. Code repositories were not. It published its incident notes and a detailed postmortem.

One-minute true story. AI-generated illustration with narration by an AI voice.

When did your team last prove it could restore a backup?

What the captions say

  1. Jan 2017 · GitLab · ~300 GB deleted
  2. Two servers · one digit apart
  3. db1 = live · db2 = copy
  4. db1
  5. ~300 GB → ~4.5 GB
  6. 5 backup methods · none working reliably
  7. "We accidentally deleted production data and might have to restore from backup."
  8. One copy · made by hand · 6 hours earlier
  9. ~18 hours · live on YouTube · ~5,000 watching
  10. Saved by the man who deleted it

Sources

More video reports