Site Reliability Engineering Technical Lead
Limerick, County Limerick, Ireland · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or Engineering or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About AMCS Group
AMCS Group is a sustainability-focused software company headquartered in Ireland with a global presence in Europe, the USA, and Australasia. Employing over 1,300 skilled professionals across 22 countries, AMCS delivers innovative SaaS solutions designed to support resource-intensive industries in achieving carbon neutrality while enhancing efficiency and profitability. Their Performance Sustainability software currently benefits over 5,000 customers worldwide, aiding them in improving both business outcomes and environmental resilience.
Company Culture
AMCS fosters a collaborative and creative work environment that values connection—to the work, customers, colleagues, and community. Retaining its roots as an Irish-founded company with a start-up mentality, AMCS offers employees a platform to develop meaningful careers while contributing to impactful global sustainability goals.
Role Overview
The company is looking for an experienced and driven DevOps/Site Reliability Engineering Technical Lead to join their engineering team. This role entails not only mentoring DevOps engineers but actively participating in architectural decisions related to infrastructure and application development. The primary focus will be on ensuring system reliability and maintaining an excellent customer experience.
Key Responsibilities
- Collaborate with development and business teams to establish Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) that accurately reflect customer experience.
- Lead incident management activities by improving incident detection, diagnosis, and resolution processes, aiming to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced alerting, tooling, and on-call procedures.
- Enhance and manage the monitoring and observability infrastructure, including Prometheus, Grafana, Mimir, Loki, Tempo, and OpenTelemetry, focusing on customer-centric operational effectiveness.
- Conduct blameless root cause analyses and postmortem reviews to convert incidents into lasting improvements that close feedback loops between development and operations teams.
- Maintain platform availability and responsiveness by identifying and addressing performance bottlenecks before they affect customers.
- Integrate Artificial Intelligence and Large Language Models into operations for incident triage, log/trace examination, automated runbook execution, and anomaly detection to shorten MTTR and reduce on-call burden.
- Optimize cloud infrastructure costs by right-sizing workloads and eliminating waste, ensuring cost-effective scaling across Azure, AWS, GCP, and container orchestration environments such as Docker and Kubernetes.
- Develop automated remediation workflows to minimize manual operational work, enabling self-healing systems that invoke human intervention only when necessary.
- Contribute to architectural design decisions ensuring alignment with organizational goals and infrastructure best practices.
Indicators of Success
- Reliable and actionable alerting systems that reduce noise and build team trust.
- Decreased frequency and severity of production incidents, with root causes effectively addressed.
- Strong, two-way communication between product teams and SRE to influence product decisions based on operational insights.
- Reduction in repetitive operational tasks through automation and self-healing capabilities.
Requirements and Qualifications
- Bachelor's degree in Computer Science, Engineering, or related discipline, or equivalent professional experience.
- Minimum of 5 years of experience in DevOps, Site Reliability Engineering, or associated areas, including at least 2 years in leadership or mentorship roles.
- Comprehensive knowledge of cloud platforms including Azure, AWS, and GCP with hands-on cloud architecture proficiency.
- Proven ability to provide architectural guidance promoting system performance and scalability.
- Extensive experience with container orchestration technologies, notably Kubernetes.
- Skilled in scripting languages such as PowerShell, Python, and Bash.
- Familiarity with monitoring and logging tools, specifically Prometheus, Grafana, and related Grafana Stack components.
- Experience with infrastructure automation tools like Ansible, Terraform, or Chef.
- Excellent leadership, communication, and collaboration capabilities.
Preferred Skills
- Experience working with CI/CD pipelines and tools like Azure DevOps, Jenkins, GitLab CI, or CircleCI.
- Certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or cloud certifications for Azure, AWS, or GCP.
- Knowledge of cloud security best practices and compliance standards.
- Familiarity with Agile project management methodologies and tools.
Minimum education
Bachelor's Degree
Industry
Software Development