Site Reliability Engineering Technical Lead
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About AMCS Group
AMCS Group is a sustainability-focused software company headquartered in Ireland with offices across Europe, the USA, and Australasia. Employing more than 1,300 professionals in 22 countries, AMCS delivers technology solutions that support a carbon-neutral future. Their innovative SaaS products help improve efficiency and sustainability within resource-intensive sectors, serving over 5,000 customers in 23 countries. They foster a connected culture emphasizing openness, collaboration, and creativity, with startup roots and a commitment to positively impacting communities.
Role Overview
The company is looking to hire an experienced DevOps/Site Reliability Engineering (SRE) Technical Lead who possesses a strong cloud computing background and a drive to enhance operational excellence. This leadership role entails mentoring DevOps engineers, participating in architectural decisions, and focusing on system reliability and exceptional customer experience. Collaborating closely with cross-functional teams, the Lead will ensure the systems' scalability, security, and high availability.
Key Responsibilities
- Develop Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) in partnership with development and business stakeholders that accurately represent customer experience.
- Lead resolution efforts during complex incidents, continuously improving detection and remediation processes, including refining alerting systems and on-call protocols to minimize detection and recovery times.
- Enhance the monitoring and observability infrastructure using tools such as Prometheus, Grafana, Mimir, Loki, Tempo, and OpenTelemetry, always with a focus on customer outcomes.
- Facilitate blameless root cause analyses and postmortems to convert incidents into sustainable operational improvements.
- Ensure system availability and responsiveness align with customer expectations by identifying and resolving performance bottlenecks proactively.
- Leverage AI and Large Language Models (LLMs) to improve incident triage, log and trace analysis, runbook execution, and anomaly detection, aiming to reduce mean time to recovery and on-call workload.
- Optimize cloud resources across platforms including Azure, AWS, and GCP, managing workload sizing and scaling for cost efficiency in containerized environments such as Docker and Kubernetes.
- Automate repetitive operational tasks to decrease toil within SRE, enabling self-healing processes where possible and escalating to human operators only when necessary.
- Contribute to architectural design and decision-making, ensuring alignment with organizational goals and best engineering practices.
Measures of Success
- High-quality alerting system that prioritizes actionable notifications and significantly reduces noise.
- Declining frequency and severity of production incidents as root causes are addressed rather than temporarily mitigated.
- Establishment of a strong feedback loop between product engineering and SRE teams to continuously improve reliability and product features.
- Significant reduction in repetitive operational work through automation, increasing time for impactful improvements.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent experience.
- Minimum of five years in DevOps, SRE, or related roles, including at least two years in a leadership or mentorship capacity.
- Extensive knowledge and practical experience with cloud platforms such as Azure, AWS, and Google Cloud Platform.
- Demonstrated ability in architectural oversight with sound judgment in achieving performance, scalability, and reliability goals.
- Proficiency in container orchestration, especially Kubernetes.
- Strong scripting skills in at least one language such as PowerShell, Python, or Bash.
- Experience with monitoring and logging tools including Prometheus, Grafana, and the Grafana stack.
- Familiarity with automation technologies like Ansible, Terraform, or Chef.
- Excellent leadership skills coupled with effective communication and teamwork abilities.
Preferred Skills
- Experience managing CI/CD pipelines using tools like Azure DevOps, Jenkins, GitLab CI, or CircleCI.
- Relevant certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or cloud certifications from Azure, AWS, or GCP.
- Understanding of security best practices and compliance considerations in cloud environments.
- Experience with Agile development methodologies and project management platforms.
Minimum education
Bachelor's Degree