A

Site Reliability Engineering Technical Lead

AMCS Group

Dublin, County Dublin, Ireland · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
1 day ago
Work mode
In office
Education
Bachelor's degree in Computer Science or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About AMCS Group

AMCS Group is a sustainability-focused software company headquartered in Ireland with offices across Europe, the USA, and Australasia. Employing more than 1,300 professionals in 22 countries, AMCS delivers technology solutions that support a carbon-neutral future. Their innovative SaaS products help improve efficiency and sustainability within resource-intensive sectors, serving over 5,000 customers in 23 countries. They foster a connected culture emphasizing openness, collaboration, and creativity, with startup roots and a commitment to positively impacting communities.

Role Overview

The company is looking to hire an experienced DevOps/Site Reliability Engineering (SRE) Technical Lead who possesses a strong cloud computing background and a drive to enhance operational excellence. This leadership role entails mentoring DevOps engineers, participating in architectural decisions, and focusing on system reliability and exceptional customer experience. Collaborating closely with cross-functional teams, the Lead will ensure the systems' scalability, security, and high availability.

Key Responsibilities

  • Develop Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) in partnership with development and business stakeholders that accurately represent customer experience.
  • Lead resolution efforts during complex incidents, continuously improving detection and remediation processes, including refining alerting systems and on-call protocols to minimize detection and recovery times.
  • Enhance the monitoring and observability infrastructure using tools such as Prometheus, Grafana, Mimir, Loki, Tempo, and OpenTelemetry, always with a focus on customer outcomes.
  • Facilitate blameless root cause analyses and postmortems to convert incidents into sustainable operational improvements.
  • Ensure system availability and responsiveness align with customer expectations by identifying and resolving performance bottlenecks proactively.
  • Leverage AI and Large Language Models (LLMs) to improve incident triage, log and trace analysis, runbook execution, and anomaly detection, aiming to reduce mean time to recovery and on-call workload.
  • Optimize cloud resources across platforms including Azure, AWS, and GCP, managing workload sizing and scaling for cost efficiency in containerized environments such as Docker and Kubernetes.
  • Automate repetitive operational tasks to decrease toil within SRE, enabling self-healing processes where possible and escalating to human operators only when necessary.
  • Contribute to architectural design and decision-making, ensuring alignment with organizational goals and best engineering practices.

Measures of Success

  • High-quality alerting system that prioritizes actionable notifications and significantly reduces noise.
  • Declining frequency and severity of production incidents as root causes are addressed rather than temporarily mitigated.
  • Establishment of a strong feedback loop between product engineering and SRE teams to continuously improve reliability and product features.
  • Significant reduction in repetitive operational work through automation, increasing time for impactful improvements.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent experience.
  • Minimum of five years in DevOps, SRE, or related roles, including at least two years in a leadership or mentorship capacity.
  • Extensive knowledge and practical experience with cloud platforms such as Azure, AWS, and Google Cloud Platform.
  • Demonstrated ability in architectural oversight with sound judgment in achieving performance, scalability, and reliability goals.
  • Proficiency in container orchestration, especially Kubernetes.
  • Strong scripting skills in at least one language such as PowerShell, Python, or Bash.
  • Experience with monitoring and logging tools including Prometheus, Grafana, and the Grafana stack.
  • Familiarity with automation technologies like Ansible, Terraform, or Chef.
  • Excellent leadership skills coupled with effective communication and teamwork abilities.

Preferred Skills

  • Experience managing CI/CD pipelines using tools like Azure DevOps, Jenkins, GitLab CI, or CircleCI.
  • Relevant certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or cloud certifications from Azure, AWS, or GCP.
  • Understanding of security best practices and compliance considerations in cloud environments.
  • Experience with Agile development methodologies and project management platforms.

Minimum education

Bachelor's Degree

Tools & software

Docker required Kubernetes required Prometheus required

How they work

Communication Teamwork & Collaboration Problem Solving Leadership
🤖
Online · instant AI help
Broxer