KAUST (King Abdullah University of Science and Technology)

Senior HPC Systems Administrator

KAUST (King Abdullah University of Science and Technology)

Makkah, Makkah Province, Saudi Arabia · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
16 часов назад
Work mode
In office
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or equivalent
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Position Overview

We seek a dedicated and experienced Senior HPC Systems Administrator to join the Supercomputing Laboratory at KAUST. The role involves overseeing a high-performance computing cluster consisting of roughly 600 CPU and GPU nodes, managing storage systems, and handling InfiniBand and Ethernet networks. This position supports daily operations and assists a wide range of researchers in computational science, engineering, big data analytics, and AI/ML workloads.

Key Duties

  • Deliver prompt and effective user assistance through various channels including phone, email, walk-in, and ticketing systems while upholding excellent customer service standards.
  • Install, configure, and maintain HPC subsystems such as compute nodes, storage solutions, and networks using configuration management tools like Ansible or Puppet.
  • Deploy and sustain cluster management and monitoring applications essential for HPC operations.
  • Manage the Slurm workload manager by administering QOS policies, user accounts, accounting, and scripting automation in Python and C++.
  • Create and update automation scripts primarily in Bash and Python to enhance system administration efficiency.
  • Implement and support container platforms (such as Singularity/Apptainer and Docker) for HPC task execution.
  • Conduct periodic performance benchmarking of CPUs, memory, network fabrics, and storage to optimize and tune hardware and software components.
  • Apply security protocols including node hardening, kernel updates, and compliance maintenance across systems.
  • Administer parallel file systems like Lustre, GPFS, Weka, or Vast, focusing on performance tuning and capacity management.
  • Collaborate directly with academic and industrial partners to support computational, engineering, data analysis, and AI/ML research activities alongside application support teams.
  • Develop internal software tools and utilities to facilitate research on cluster infrastructure.
  • Lead proof-of-concept initiatives, evaluate new technologies, and promote system innovations.
  • Coordinate with external vendors and service providers to troubleshoot and resolve hardware/software issues efficiently.
  • Create and maintain comprehensive user guides, operational procedures, and training content in the internal documentation repository.
  • Maintain up-to-date knowledge of HPC advancements, participate in relevant conferences, and spearhead benchmarking efforts to guide future infrastructure investments.

Required Competencies

  • Extensive experience supporting users in computational science, engineering, data analytics, and AI applications in HPC contexts.
  • Thorough knowledge of Linux system administration (RHEL, Rocky Linux, CentOS) within large-scale HPC environments.
  • Proficient in HPC programming languages and models including Fortran, C/C++, Python, MPI, OpenMP, CUDA, and OpenACC.
  • Proven background managing intricate HPC infrastructure including parallel file systems, job schedulers, high-speed networks, and monitoring tools.
  • Hands-on experience with automation and configuration management tools such as Ansible or Puppet.
  • Understanding of project management methodologies and proven ability to support research activities collaboratively.
  • Strong analytical thinking, problem-solving aptitude, and effective decision-making capabilities.
  • Self-driven with initiative to identify improvements and follow through to completion.
  • Ability to manage multiple projects simultaneously, meeting deadlines with quality deliverables.
  • Effective collaboration skills across diverse teams including researchers, application teams, and vendors within international multicultural settings.
  • Excellent verbal and written English communication skills for technical reporting and presentations.

Qualifications and Experience

  • Minimum education requirement: Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field.
  • At least five years of experience in large-scale computing platform support and subsystem administration.
  • Proficient in troubleshooting complex hardware problems and documenting detailed root cause analyses.
  • Experienced managing parallel storage systems such as Lustre, GPFS, Weka, Vast, or equivalents.
  • Skilled in performance benchmarking of CPU, memory, network, and storage components in HPC setups.
  • Experienced administering workload managers like Slurm, LSF, or PBS.
  • Strong Linux system administration skills specific to RHEL, Rocky Linux, or CentOS environments.
  • Familiar with configuration management practices using tools like Ansible or Puppet.
  • Demonstrated capacity to collaborate with researchers, support teams, and vendors to resolve difficult issues.
  • Keen ability to drive initiatives to successful execution with cross-functional collaboration.
  • Knowledge of Kubernetes and container orchestration platforms is advantageous.

Minimum education

Master's Degree

How they work

Communication Teamwork & Collaboration Problem Solving Time Management Initiative

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer