Inference Performance Engineer
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 17 hours ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
As an Inference Performance Engineer, you will be responsible for the efficiency and effectiveness of our model inference systems. Your efforts will directly influence how our platform handles workloads, traffic patterns, and hardware variations to maintain optimal performance at minimal cost.
Key Responsibilities
- Enhance throughput, cost-efficiency, and reduce latency by managing KV-cache, implementing continuous batching, speculative decoding, and quantization techniques.
- Optimize workloads related to long-context prefill and decoding based on live production data.
- Balance routing across infrastructure and external providers by considering cost, capacity, and performance metrics.
- Engage with serving engines such as vLLM, SGLang, and TensorRT-LLM, performing kernel-level modifications when necessary.
- Develop profiling and monitoring tools to identify resource usage in time, memory, and compute.
Qualifications
- A minimum of 5 years' experience in machine learning systems, inference infrastructure, or performance engineering demonstrating quantifiable improvements in cost or latency.
- Profound knowledge of model serving including prefill, decoding, memory bandwidth management, batching, and concurrency control.
- Hands-on experience with serving engines like vLLM, SGLang, or TensorRT-LLM in production environments.
- Strong programming expertise in Python and competency in C++, Rust, or other systems programming languages.
- Experience in GPU performance optimization including CUDA, NCCL, mixed precision computing, efficient memory layouts, kernel development, and quantization methods.
Team and Culture
We value collaborative team members who can lighten the workload with positivity and creativity. Although perfection is not expected, adaptability and courage to pursue innovative ideas are essential. We warmly welcome applicants who may not fulfill every qualification but are eager to contribute.
Company Overview
Our mission is to create AI systems that constantly evolve and adapt in real time, making intelligence more flexible, personalized, and universally available. We emphasize efficiency as the foundation for broad access and innovative impact, championing a high-talent environment to pioneer the future of adaptive AI.
Benefits
- Flexible work arrangements: While based in Dublin, the company offers in-person collaboration and global team offsites to foster connectivity.
- Annual travel stipend (Adaption Passport) to encourage exploration of new countries and broaden personal horizons alongside professional growth.
- Weekly meal allowance supporting takeout or grocery delivery to maintain convenience and wellness.
- Comprehensive health coverage and generous paid leave to support overall well-being.