Mei Tanaka
Site Reliability Engineer
Portland, OR | mei.tanaka@example.com | linkedin.com/in/mei-tanaka-example
Summary
Site reliability engineer with six years running production services on Linux and Kubernetes. Focused on setting sensible service level objectives, cutting repetitive operational work, and turning incidents into changes that stop them recurring.
Experience
Site Reliability Engineer - Video streaming company, Portland2021 - Present
- Defined SLOs and error budgets for six customer-facing services with product teams, which ended a recurring argument about release pace and cut unplanned outages by 40% year over year.
- Reduced on-call pages from about 60 to 22 a month by tuning noisy alerts and automating three recurring remediations in Python, improving the rotation for a team of eight.
- Led incident response for a database failover that caused 26 minutes of degraded service, and wrote the blameless postmortem whose five follow-ups were all closed within the quarter.
- Built runbooks and a capacity model for peak traffic events, which carried the platform through its largest launch with no capacity-related incident.
Systems Engineer - Managed hosting provider, Portland2018 - 2021
- Managed 400+ Linux servers, replacing manual patching with configuration-managed rollouts and cutting patch time from days to hours.
- Introduced monitoring and alerting where none existed, giving support the first view of customer-affecting issues before tickets arrived.
- Migrated 30 services to containers on Kubernetes with no customer-visible downtime.
Skills
Core: Linux, Python, Go, Monitoring & alerting, Incident management, SLO/SLA, Kubernetes, On-call
Practices: Error budgets, Runbooks, Capacity planning, Distributed tracing, Chaos engineering
Education
B.S. Computer Science, State University, 2018