NVIDIA

Senior GPU Networking Architect

NVIDIA

Germany · Full Time

Be the first to apply

Experience
5+ yrs
Salary
PLN 292,500 – PLN 650,000 / year
Openings
1
Posted
4 days ago
Work mode
In office
Education
M.Sc. or equivalent
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About NVIDIA

NVIDIA has been a pioneer in the realms of computer graphics, PC gaming, and accelerated computing for over 25 years. Today, the company is innovating at the forefront of AI, shaping the future of computing with GPUs that drive intelligent systems including robotics and autonomous vehicles. As an NVIDIA employee, you will be part of a diverse, supportive environment that inspires excellence and innovation.

Role Overview

We are seeking a Senior GPU Networking Architect to join our networking software team. This position requires robust expertise in GPU architecture and programming to design and enhance GPU communication kernels that are fundamental to large-scale AI platforms. The role bridges GPU computing with networking by integrating communication primitives closely aligned with GPU hardware capabilities.

Key Responsibilities

  • Design, implement, and optimize GPU communication kernels facilitating collective and point-to-point operations in expansive AI systems.
  • Utilize in-depth GPU architectural knowledge—including thread scheduling, memory hierarchy, and execution pipelines—to enhance kernel efficiency, reduce latency, and enable computation-communication overlap.
  • Create GPU-resident communication primitives and device-level APIs to support fine-grained, kernel-driven data transfer across nodes and accelerators.
  • Conduct end-to-end profiling and tuning of GPU kernels to identify and resolve performance bottlenecks at the intersection of compute, memory, and networking.
  • Collaborate closely with teams in network software, hardware, and AI frameworks to develop communication strategies aligned with GPU execution and evolving AI model structures.
  • Develop proofs-of-concept, run experiments, and quantitatively analyze new communication strategies to validate their readiness for production deployment.
  • Contribute to advancing programming models that expose GPU-aware networking features to developers.

Required Qualifications

  • Minimum 5 years of practical experience programming with CUDA, especially in writing and optimizing complex GPU kernels.
  • A Master’s degree or equivalent in computer science, computer engineering, or related disciplines.
  • Strong grasp of GPU architecture principles, including warp scheduling, shared memory management, L2 cache utilization, memory coalescing, occupancy tuning, and asynchronous execution mechanisms.
  • Proficient in systems-level C/C++ programming within performance-sensitive contexts.
  • Knowledge of GPU data movement methods such as GPUDirect RDMA and GPU-initiated communication processes.
  • Skilled in interpreting GPU performance analytics tools like Nsight Compute and Nsight Systems to derive actionable kernel optimizations.
  • Excellent teamwork and communication capabilities in an international, interdisciplinary setting.

Preferred Skills

  • Experience in designing or optimizing communication kernels within frameworks like NCCL, NVSHMEM, or equivalent.
  • Understanding of distributed deep learning parallelism schemes (data, tensor, pipeline, expert, mixture-of-experts) and their influence on GPU communication patterns.
  • Experience with RDMA, InfiniBand, high-speed networking, and GPU system topologies (NVLink, NVSwitch, PCIe), and their effects on communication kernel architecture.
  • Familiarity with overlapping techniques such as kernel pipelining, persistent kernels, or cooperative groups for latency mitigation.
  • Proven background optimizing large-scale large language model training or inference workflows, with hands-on exposure to frameworks like PyTorch, TensorRT-LLM, or vLLM and knowledge of advanced serving architectures including disaggregated serving.

Additional Information

NVIDIA ranks among the most sought-after employers worldwide, offering competitive compensation packages and comprehensive benefits. Salary bands are determined based on location, experience, and position seniority; specifically, in Poland the base salary ranges from 292,500 to 507,000 PLN for Level 4 roles, and 375,000 to 650,000 PLN for Level 5 positions.

For a detailed overview of benefits for you and your family, NVIDIA provides comprehensive resources reflecting their commitment to employee wellbeing.

Minimum education

Master's Degree

Industry

Semiconductors

How they work

Teamwork & Collaboration Problem Solving Creativity
🤖
Online · instant AI help
Broxer