R

Principal Software Engineer - Machine Learning Platform

Riot Games

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
SGD 253,300 – SGD 430,700 / year
Openings
1
Posted
1 دن قبل
Work mode
In office
Education
Bachelor’s degree in Computer Science or related field
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Riot Games and Role Overview

Founded in 2006 by gamer entrepreneurs, Riot Games focuses on creating player-centric games, exemplified by the acclaimed and widely played League of Legends. The AI Efficiency team builds scalable platforms and tooling to empower safe, effective use of AI to enhance work processes across the company.

As Principal Platform Engineer on this team, you will architect and maintain internal platforms, automation, and safeguards that streamline AI service development, deployment, and operation. Collaborating across engineering and infrastructure teams, you will enhance developer experiences, operational robustness, security, and observability for a growing suite of AI and developer tools.

Key Responsibilities

  • Design, develop, and evolve internal platform features that simplify building and operating AI services.
  • Create self-service workflows, reusable platform abstractions, and standardized golden paths to boost developer productivity while ensuring reliability and security.
  • Enhance platform reliability through advanced monitoring, alerting, observability, release safety, and incident preparedness.
  • Define and put into practice service health metrics such as SLIs and SLOs to balance reliability, velocity, and cost effectively.
  • Develop automation to reduce operational burden and accelerate incident detection and response.
  • Collaborate throughout the software lifecycle to embed operational readiness and maintainability in system designs.
  • Optimize CI/CD pipelines, developer workflows, and release processes for safer, faster, and repeatable software delivery.
  • Identify risks related to distributed systems, infrastructure, dependencies, and operations, driving lasting improvements.
  • Troubleshoot AI model-serving across various frameworks, runtimes, hardware, including GPU platform compatibility and model format conversion.
  • Conduct resilience, failure-mode, and recovery testing to reveal systemic weaknesses before impacting users.
  • Evaluate and integrate AI-driven engineering tools that enhance code quality, security, performance, and productivity.
  • Build automation pipelines combining conventional CI/CD with agentic workflows like automated code review, bug detection, regression testing, and remediation.
  • Partner with engineers in safe, auditable, and measurable deployment of AI agents for pull request reviews, diagnostics, UI/UX validation, accessibility, and production readiness checks.
  • Define governance protocols including guardrails, approval workflows, observability, reporting, and escalation to maintain trustworthy AI-assisted automation.
  • Establish evaluation frameworks measuring quality improvements, false positives, latency, cost, risks, and impact on engineering output for AI-native tools.
  • Lead or support incident response and systemic improvements for critical internal platforms with focus on thorough resolution.
  • Advocate platform and operational excellence through documentation, runbooks, standards, and tools that elevate engineering practices company-wide.

Essential Qualifications

  • Bachelor’s degree in Computer Science or related field, or equivalent experience.
  • More than 5 years’ experience in Platform Engineering, Infrastructure, Site Reliability Engineering, DevOps, or Developer Experience roles supporting production systems and workflows.
  • Expertise in programming and automation with Python, Go, or JavaScript/TypeScript.
  • Experience designing, building, or operating internal platforms, developer tooling, CI/CD, or shared infrastructure used by multiple engineering teams.
  • Operational experience with cloud-based production environments like AWS, GCP, or Azure.
  • Strong grasp of observability best practices including metrics, logging, tracing, dashboarding, and alert design.
  • Proven track record improving reliability and operability for distributed systems, service-oriented architectures, APIs, or platform infrastructure.
  • Incident management skills, performing root cause analysis and implementing long-term operational fixes after production issues.
  • Knowledge of container technologies and orchestration platforms such as Kubernetes or ECS.
  • Effective collaborator able to influence technical direction and communicate clearly across technical and non-technical audiences.

Preferred Competencies

  • Experience supporting AI/ML platforms, inference systems, model serving, data pipelines, or GPU-accelerated workloads.
  • Familiarity with defining and leveraging SLOs, error budgets, and reliability metrics for prioritization and engineering decisions.
  • The ability to apply platform product thinking such as designing self-service and golden path experiences for internal teams.
  • Proven success improving developer platforms, internal tools, or enterprise services for reliability and usability.
  • Experience with infrastructure as code tools like Terraform or Pulumi.
  • Knowledge of security, access control, secrets management, and operational hardening in production.
  • Ability to balance availability, latency, cost, and usability in large-scale systems.
  • Experience mentoring engineers and elevating platform and reliability standards through leadership.
  • Familiarity with AI-assisted software engineering tools for code review, static analysis, test generation, and operational automation.
  • Insight into emerging AI agent workflows interacting with source control, CI/CD, browser automation, and developer platforms.
  • Competence with browser automation and testing frameworks such as Playwright for UI validation, accessibility, regression tests, and workflow automation.
  • Experience enforcing governance, review loops, and quality controls for automated and AI-assisted engineering systems.
  • Comfort operating at the crossroads of platform engineering, reliability, developer productivity, and AI-native software delivery approaches.

Additional Information and Perks

  • Extensive relocation support offered.
  • Comprehensive health coverage including yourself, partner, and children.
  • Flexible paid time off policy.
  • Retirement plan with company matching benefits.
  • Life insurance, parental leave, short- and long-term disability coverage.
  • Play Fund to deepen understanding of players and community through gaming experiences.
  • Matching donations and support for nonprofit engagements.

Compensation

The salary range for this position based in Singapore is between SGD 253,300 and SGD 430,700 per annum.

Level

Lead

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required

How they work

Communication Teamwork & Collaboration Problem Solving Leadership Initiative

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer