- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 6 minutes ago
- Work mode
- In office
- Education
- Bachelor's degree or equivalent
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We seek a skilled LLM Inference Performance Engineer dedicated to enhancing the performance of large language model (LLM) inference across TPUs, GPUs, and other AI acceleration hardware. The role focuses on boosting inference speed, scalability, and efficiency by working on inference engines, kernels, compilers, and runtime systems.
Key Responsibilities
- Enhance LLM inference capabilities on TPU, GPU, and alternative AI acceleration platforms.
- Design and refine inference backends, computational kernels, and runtime modules.
- Optimize critical workloads such as Attention mechanisms, GEMM operations, KV Cache management, Sampling processes, and fused kernel computations.
- Employ technologies including JAX, XLA, Pallas, CUDA, Triton, or similar frameworks to drive performance improvements.
- Create benchmarking and profiling utilities to detect and analyze performance bottlenecks.
- Work collaboratively with teams specializing in model development, inference, compilers, and hardware to elevate production-level performance.
Qualifications and Requirements
- Possess a bachelor's degree or equivalent expertise in computer science, engineering, machine learning, systems, or related disciplines.
- Experience in areas such as LLM inference, machine learning systems, hardware acceleration, or performance tuning.
- Proficiency in programming with C++ and/or Python is essential.
- Knowledge of GPU and TPU architectures, memory systems, kernel design, and performance considerations for ML workloads.
- Hands-on experience with performance profiling and benchmarking methodologies.
Minimum education
Bachelor's Degree
Tools & software
Python
required
C
required