CoreWeave

Staff Software Engineer, Inference

CoreWeave

Singapore · Full Time

Be the first to apply

Experience
8–12 yrs
Salary
Openings
1
Posted
6 天前
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About CoreWeave

CoreWeave is a leading cloud platform purpose-built to enable AI innovation. Founded in 2017 and becoming publicly traded in March 2025, CoreWeave offers high-performance infrastructure combined with expert technical support, empowering leading AI labs, startups, and enterprises worldwide to accelerate breakthroughs with confidence.

Role Overview

As a Staff Software Engineer on the Inference Platform team based in Singapore, you will provide technical leadership and drive architectural decisions for CoreWeave's Kubernetes-native inference platform. This platform operates at massive scale, supporting AI workloads that demand low latency and high throughput. You will focus on system-wide improvements involving request routing, GPU resource management, scheduling, and optimization to enhance performance, efficiency, and reliability of real-time inference services.

Key Responsibilities

  • Lead the design and implementation of cross-team initiatives to improve inference platform performance and reliability.
  • Optimize latency, throughput, and GPU utilization for distributed inference workloads.
  • Work extensively with Kubernetes infrastructure, focusing on scheduling, batching, and memory usage enhancements.
  • Influence and set engineering directions that impact multiple teams and services.

Candidate Profile

  • 8 to 12+ years of experience architecting and managing large-scale distributed systems or cloud platforms.
  • Proven ability to lead cross-organizational technical projects.
  • Proficient in programming languages such as Go, Python, or C++.
  • Expertise in Kubernetes production environments, including orchestration and service design.
  • Strong knowledge of distributed systems principles, networking, and systems performance tuning.
  • Hands-on experience building low-latency, high-throughput systems meeting strict percentile latency targets (P95/P99).
  • Familiarity with inference systems requiring batching, caching, and memory optimization strategies.
  • Skilled in metrics-driven performance improvements and system monitoring.
  • Understanding of mixed precision formats (e.g., BF16, FP8) and streaming inference workloads.
  • Preferred experience with inference frameworks like vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe.
  • Background in GPU system optimization technologies such as CUDA, NCCL, RDMA, NUMA, and GPU interconnects.
  • Experience in leading initiatives spanning multiple teams or departments.
  • Exposure to large-scale AI/ML infrastructure or hyperscale cloud ecosystems.

Personal Qualities

  • Passionate about designing and enhancing high-scale distributed systems.
  • Keen interest in AI inference, GPU technologies, and innovative performance optimization methods.
  • Expert in building fault-tolerant, low-latency platforms and promoting organization-wide system improvements.

Company Culture and Benefits

CoreWeave offers a dynamic and fast-paced work environment emphasizing learning, collaboration, and ownership. The company fosters innovation and empowers employees while focusing on delivering excellent client experiences. Benefits include competitive base salary, equity, flexible vacation policies, and comprehensive health coverage. Compensation is set based on individual skills, experience, and market standards, ensuring fairness and internal equity.

Level

Mid

How they work

Teamwork & Collaboration Leadership Learning Agility

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer