adaption

Inference Performance Engineer

adaption

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
15 hours ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

As an Inference Performance Engineer, you will take full ownership of the cost efficiency and operational performance of our model inference stack. Your efforts will be critical in maintaining high throughput and low latency while adapting to changing workloads, traffic patterns, and hardware configurations.

Key Responsibilities

  • Enhance throughput, reduce costs, and improve latency by managing key-value caching, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode processes based on live production data.
  • Balance routing decisions across internal infrastructure and third-party providers according to factors like cost, capacity, and performance.
  • Engage deeply with serving engines such as vLLM, SGLang, and TensorRT-LLM, including modifications at levels beneath the framework as necessary.
  • Develop profiling and measurement tools to analyze time, memory, and computational resource consumption.

Candidate Qualifications

  • At minimum five years of experience in machine learning systems, inference infrastructure, or performance optimization with proven results in lowering latency or cost.
  • Strong understanding of model serving pipelines, including prefill and decode stages, efficient memory bandwidth usage, batching, and concurrency control.
  • Hands-on production experience with inference serving tools like vLLM, SGLang, or TensorRT-LLM.
  • Advanced programming skills in Python alongside proficiency in at least one systems programming language such as C++ or Rust.
  • Deep expertise in GPU performance optimization techniques, including use of CUDA, NCCL, mixed precision arithmetic, memory layouts, kernel functions, and quantization.

About Our Company

We focus on developing AI systems capable of real-time adaptation and evolution, making AI more flexible, personalized, and accessible. By emphasizing efficiency, we aim to democratize innovation, ensuring breakthroughs serve a broad audience. Our culture values high talent density and creative problem-solving to push the boundaries of continuous adaptation.

Employee Benefits

  • Flexible working arrangements: collaborative in-person work in the Bay Area, global-distributed team structure, and regular team retreats.
  • Annual travel stipend termed the Adaption Passport to encourage exploration of new countries.
  • Weekly lunch stipend applicable to takeout or grocery deliveries.
  • Comprehensive health care coverage combined with ample paid time off.

How they work

Teamwork & Collaboration Problem Solving Adaptability Creativity
🤖
Online · instant AI help
Broxer