Epam Systems

Lead Kernel Engineer/Architect

Epam Systems

Berlin, Germany (Hybrid) · Full Time

Be the first to apply

Experience
12+ yrs
Salary
Openings
1
Posted
1 day ago
Work mode
Hybrid
Education
Bachelor's degree or equivalent
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We invite an experienced Lead Kernel Engineer/Architect to join our German team operating in a hybrid setup. The position focuses on pushing the boundaries of advanced hardware accelerators to enhance AI performance and scalability. You will spearhead the refinement of pivotal machine learning operations used in extensive training and inference, working with the latest hardware like TPUs and GPUs, along with advanced ML models and performance toolkits. Your contributions will facilitate faster AI research progress and deployment on cloud platforms and within open-source communities.

Key Responsibilities

  • Architect and optimize performant kernels tailored for TPU and GPU platforms leveraging low-level programming tools such as Pallas, Triton, or Mosaic.
  • Develop and sustain performance-focused infrastructure including benchmarking suites, autotuning mechanisms, regression testing frameworks, and related tooling.
  • Partner closely with machine learning framework developers (e.g., JAX, PyTorch) and compiler engineers (XLA/MLIR) to embed custom kernels and mitigate performance bottlenecks.
  • Monitor emerging trends in accelerator hardware, compiler technology, and AI model structures to spot areas ripe for kernel-level enhancements.
  • Create comprehensive documentation, APIs, and open-source software elements that elevate developer experience and broaden adoption.
  • Diagnose and solve intricate performance challenges impacting large-scale distributed AI training and inference systems.

Candidate Profile

  • Possess a Bachelor's degree or equivalent practical knowledge.
  • Bring over 12 years of professional experience in software or systems programming.
  • Have at least 5 years experience in software development using C++ or Python.
  • Hold 3+ years experience in software product lifecycle activities including testing, maintenance, or launch, with a minimum of 1 year in software design or system architecture.
  • Demonstrate hands-on expertise in kernel-level performance optimization for accelerators or high-performance computing environments.

Desirable Skills

  • Expertise in low-level accelerator programming (CUDA, Triton, Pallas).
  • Experience with ML frameworks like JAX or PyTorch and knowledge of optimization methods such as attention layers, Mixture of Experts (MoE), and precision tuning.
  • Deep understanding of modern hardware accelerators encompassing pipelining, data transfer, and heterogeneous computing.
  • Familiarity with compiler theory and intermediate representations (e.g., MLIR, OpenXLA).
  • Background in constructing open-source developer tools, APIs, and performance-critical libraries.
  • Strong analytical skills with an ability to thrive in multidisciplinary engineering teams.

Benefits and Perks

  • Entitlement to 30 days of annual leave.
  • Access to a company pension scheme.
  • Regular performance evaluations and feedback.
  • Discounted Fitness-First Black membership.
  • Corporate benefits program through bitkom.
  • Potential participation in an Employee Stock Purchase Plan (subject to eligibility).
  • Opportunities for learning and professional development, including internal training, coaching, certifications, and specialized courses.
  • A collaborative and friendly team environment.
  • Frequent social and corporate events.
  • Flexible working policies, including remote options.
  • Workplace recognized with various awards for culture and innovation, including Great Place To Work® certification (2026) and Kununu Top Company distinction (2022–2026).

Level

Lead

Minimum education

Bachelor's Degree

How they work

Teamwork & Collaboration Problem Solving

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer