Russell Tobin

Java Site Reliability Engineer (SRE)

Russell Tobin

Singapore · Full Time

Be the first to apply

Experience
8–15 yrs
Salary
Openings
1
Posted
1 day ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

We are seeking an experienced Java Site Reliability Engineer to join our team in Singapore. This role requires 8 to 15 years of overall IT experience, including a minimum of 5 years dedicated to SRE, Production Support, Platform Engineering, or DevOps. The candidate will have expertise in Java/J2EE and a strong background in supporting mission-critical banking or financial systems within a 24/7 production environment adhering to rigorous SLA and SLO standards.

Key Responsibilities and Skills

  • Manage and support Java-based applications with a focus on stability and performance in mission-critical financial platforms.
  • Work extensively with container technologies such as Kubernetes, OpenShift, and Docker for application deployment and orchestration.
  • Implement and maintain Infrastructure as Code (IaC) solutions using tools like Terraform, Ansible, and CloudFormation.
  • Conduct capacity planning and performance engineering to ensure optimal resource utilization and system responsiveness.
  • Lead Root Cause Analysis (RCA) and manage problem resolution processes effectively.
  • Design and execute chaos engineering practices and resilience testing to improve system reliability.
  • Utilize monitoring and logging platforms such as Prometheus, Grafana, ELK stack, and Splunk to ensure observability and alerting.
  • Employ Application Performance Monitoring (APM) tools including Dynatrace, AppDynamics, or New Relic to track and optimize system health.
  • Drive automation for operational tasks and develop self-healing capabilities for the platform.
  • Leverage CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, ArgoCD, and SonarQube to facilitate continuous integration and deployment processes.
  • Possess a working knowledge of Kafka, Hadoop/Cloudera ecosystem, and Spark to support related data processing workflows.

Tools & software

Java Docker required Kubernetes required Ansible required Terraform required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer