Professional.me

Principal Site Reliability Engineer (SRE)

Professional.me

Abu Dhabi Emirate, United Arab Emirates · Full Time

Be the first to apply

Experience
12+ yrs
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Company

Alpheya is an Abu Dhabi–based wealth-management technology firm serving banking institutions. They offer a customizable investment platform delivered as SaaS on Microsoft Azure, complemented by on-premises solutions for banks requiring them. Their platform enables a full order-to-custody lifecycle, accessible through mobile, advisor, and back-office portals.

Role Overview

This leadership role reports to the CTO and holds responsibility for ensuring seamless SaaS operations across multiple banking clients. The role entails managing service-level agreements (SLAs), incident handling, disaster recovery initiatives, cost analysis per tenant, and facilitating onboarding processes to evolve from project-based to checklist-based deployment. The aim is to enhance operational efficiency and reliability as the platform rolls out to additional banks.

Key Responsibilities

  • Lead service management for each SaaS customer, including defining SLAs, incident communication and reviews, security questionnaire responses, audit management, and maintaining effective daily support coordination with client service desks.
  • Oversee release and deployment procedures, manage the release schedule for the tenant environment, enforce change management protocols such as environment promotion and rollback, and align deployments with each bank's change and freeze periods.
  • Direct the reliability program comprising disaster recovery testing, business continuity planning, consistent backup verification, patch management, vulnerability assessments, capacity forecasting, and developing a tenant cost model for financial review.
  • Create and maintain a tenant onboarding runbook to ensure predictable and replicable new bank go-live processes.
  • Steer the European expansion by establishing operations, including data residency, disaster recovery, and support coverage in new Azure regions.
  • Manage compliance activities related to outsourcing arrangements and collaborate with bank risk teams under compliance frameworks like ISO 27001, SOC 2, and appropriate regional outsourcing regulations.

First Six Months Deliverables

  • Enable successful go-lives for four banks across two regions, ensuring clearly established SLAs, escalation processes, and incident response workflows from the outset.
  • Coordinate and document at least one comprehensive disaster recovery exercise for a production environment.
  • Implement and manage a sustainable on-call rotation spanning both operational regions without requiring extraordinary efforts.
  • Develop and present a detailed tenant cost and capacity plan encompassing the entire SaaS environment.

Candidate Profile

  • Over 12 years of hands-on experience in production operations, including leadership roles managing multi-tenant SaaS platforms.
  • Proven track record of owning customer-centered service management, including leading severity-one incidents, communicating post-incident findings to executives, and navigating customer audits.
  • Experience in building or scaling SRE or platform operations teams, including designing cross-regional on-call systems.
  • Strong technical expertise in Kubernetes and cloud environments sufficient to critically assess system failure modes, disaster recovery designs, and capacity-related assertions.
  • Operational knowledge of Microsoft Azure components such as regions, availability zones, networking, private connectivity, identity management, and quota planning, including handling negotiations with Microsoft and banking security teams.
  • Familiarity with certification and audit standards including ISO 27001, SOC 2, and bank outsourcing regulations applicable to Europe or the Gulf.

Desirable but Not Required

  • Experience in wealth management, brokerage, or capital markets domains.
  • Hands-on involvement with streaming or workflow management tools such as Temporal or Kafka.
  • Background with on-premises software deployment tailored for enterprise clients.

Compensation and Reporting

The role offers a competitive remuneration package, a senior-level position reporting directly to the CTO, and the opportunity to manage a platform transitioning from early customers toward a robust production environment.

Additional Information

By applying, candidates consent to their CV being processed and retained for this and future employment considerations.

Level

Lead

Industry

FinTech

Tools & software

Kubernetes required Microsoft Azure required

How they work

Communication Teamwork & Collaboration Problem Solving Leadership Accountability
🤖
Online · instant AI help
Broxer