Cloud & Platform
7 posts in this category.
What we got wrong about multi-region
eu-de-1 has been live for five years. Some of the assumptions we built it on turned out to be wrong.
Read post →What a real post-incident review looks like
We publish a review for every incident with customer impact. Here is the structure and why blame is absent from it.
Read post →How we choose which alerts wake someone up
Our page volume dropped 62% and our incident detection got faster. Those are related.
Read post →Scheduled maintenance is not an incident
How we schedule, announce and run disruptive changes — and why we announce ones that turn out to need no downtime.
Read post →Adding a second data centre without a second team
FRA2 went live last year. The infrastructure was straightforward; keeping the operational model singular was not.
Read post →Moving 2.4 petabytes without anybody noticing
We consolidated four regional archives into one object store over eleven weeks. Nobody filed a ticket.
Read post →ORBIS Cloud is live, and we moved our own workloads first
The managed service launches today in eu-nl-1. We have been running our own processing on it since January.
Read post →