Senior Site Reliability Engineer
Melbourne, Victoria, Australia · Contract
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are seeking a Site Reliability Engineer specializing in observability to ensure critical systems for clients' stores, warehouses, and digital channels operate effectively during peak periods. This position involves ownership of the Prometheus and Grafana monitoring stack to convert telemetry data from checkout, payment, inventory, and store systems into actionable dashboards, alerts, and service level objectives (SLOs), ensuring uninterrupted business operations during high-demand events like Black Friday and Christmas.
Key Responsibilities
- Develop and sustain the Prometheus and Grafana observability infrastructure deployed across Kubernetes, cloud platforms, and store and warehouse systems.
- Establish key performance indicators including golden signals, SLOs, and error budgets for checkout, payments, inventory, and store connectivity components.
- Create effective, low-noise alerting mechanisms to reduce mean time to detection and resolution (MTTD/MTTR).
- Implement tracing strategies utilizing OpenTelemetry to cover the entire customer journey from browsing to payment and stock updates.
- Lead preparedness measures for peak trading periods, including capacity planning, load testing, and real-time monitoring during events such as Black Friday, Christmas, and flash sales.
- Conduct blameless post-incident reviews and ensure every incident results in a permanent corrective action.
- Automate routine platform maintenance and incident remediation wherever feasible.
Required Qualifications and Experience
- Proven practical experience managing Prometheus and Grafana in production environments, preferably at scale.
- Familiarity with Kubernetes and cloud infrastructures such as Azure, AWS, or Google Cloud Platform.
- Working knowledge of OpenTelemetry, time series databases (TSDBs), and open metrics standards.
- Competency with Terraform, continuous integration/continuous delivery pipelines, Git version control, and scripting languages like Bash, Python, or PowerShell.
- Understanding of DevOps Research and Assessment (DORA) metrics with a track record of leveraging them for system improvements.
- Experience in retail, e-commerce, or industries characterized by seasonal demand spikes is advantageous.
Organizational Culture
Our client promotes a culture where incidents are seen as opportunities to learn rather than assign blame. Reliability is a shared responsibility across all teams, not solely the SRE group. They prioritize developing user-friendly tools and ensure on-call duties are distributed and manageable, emphasizing thorough planning during off-peak periods to maintain stability during critical business times.
How to Apply
Interested candidates are encouraged to submit their CV along with a brief description of an observability issue they have successfully addressed.
Level
Senior