i

Senior Site Reliability Engineer

iterate

Melbourne, Victoria, Australia · Contract

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 day ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

We are seeking a Site Reliability Engineer specializing in observability to ensure critical systems for clients' stores, warehouses, and digital channels operate effectively during peak periods. This position involves ownership of the Prometheus and Grafana monitoring stack to convert telemetry data from checkout, payment, inventory, and store systems into actionable dashboards, alerts, and service level objectives (SLOs), ensuring uninterrupted business operations during high-demand events like Black Friday and Christmas.

Key Responsibilities

  • Develop and sustain the Prometheus and Grafana observability infrastructure deployed across Kubernetes, cloud platforms, and store and warehouse systems.
  • Establish key performance indicators including golden signals, SLOs, and error budgets for checkout, payments, inventory, and store connectivity components.
  • Create effective, low-noise alerting mechanisms to reduce mean time to detection and resolution (MTTD/MTTR).
  • Implement tracing strategies utilizing OpenTelemetry to cover the entire customer journey from browsing to payment and stock updates.
  • Lead preparedness measures for peak trading periods, including capacity planning, load testing, and real-time monitoring during events such as Black Friday, Christmas, and flash sales.
  • Conduct blameless post-incident reviews and ensure every incident results in a permanent corrective action.
  • Automate routine platform maintenance and incident remediation wherever feasible.

Required Qualifications and Experience

  • Proven practical experience managing Prometheus and Grafana in production environments, preferably at scale.
  • Familiarity with Kubernetes and cloud infrastructures such as Azure, AWS, or Google Cloud Platform.
  • Working knowledge of OpenTelemetry, time series databases (TSDBs), and open metrics standards.
  • Competency with Terraform, continuous integration/continuous delivery pipelines, Git version control, and scripting languages like Bash, Python, or PowerShell.
  • Understanding of DevOps Research and Assessment (DORA) metrics with a track record of leveraging them for system improvements.
  • Experience in retail, e-commerce, or industries characterized by seasonal demand spikes is advantageous.

Organizational Culture

Our client promotes a culture where incidents are seen as opportunities to learn rather than assign blame. Reliability is a shared responsibility across all teams, not solely the SRE group. They prioritize developing user-friendly tools and ensure on-call duties are distributed and manageable, emphasizing thorough planning during off-peak periods to maintain stability during critical business times.

How to Apply

Interested candidates are encouraged to submit their CV along with a brief description of an observability issue they have successfully addressed.

Level

Senior

Tools & software

Python required

How they work

Teamwork & Collaboration Problem Solving Attention to Detail Accountability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer