Handling Outages with Confidence: Incident Management for Websites and Platforms
What happens when a platform with millions of SIM cards worldwide goes down? At the TYPO3 Developer Days, Steffen Gebert shared insights into how to handle critical outages. His presentation offered a practical look at how alerting, on-call teams, and follow-up work together to ensure that issues are resolved quickly and professionally.
A key point: Alerts should be based on the actual impact on customers rather than on individual metrics such as CPU utilization. This is the only way to prevent important warnings from getting lost in a flood of meaningless messages.
For organizing on-call shifts, Gebert recommended specialized tools that aggregate alerts from various sources and reliably deliver them via app, phone call, or text message—even outside of regular working hours. He emphasized the importance of having a team of at least three to four people, clear agreements on response times, and fair compensation for on-call duty.
In the event of major disruptions, the role of an incident manager has proven effective in his company. This person guides the team through the crisis, maintains an overview, and keeps the team calm while others focus on the technical solution. His advice: First limit the damage; the detailed root cause analysis comes later.
Finally, the discussion turned to postmortems: structured follow-ups with a clear timeline, concrete measures, and a focus on human and organizational aspects as well—not just on technology. In summary: Those who improve processes step by step and value their team are better prepared for the next outage.
This page contains automatically translated content.
