Senior HPC Systems Administrator
KAUST (King Abdullah University of Science and Technology)
Makkah, Makkah Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 16 часов назад
- Work mode
- In office
- Education
- Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or equivalent
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Position Overview
We seek a dedicated and experienced Senior HPC Systems Administrator to join the Supercomputing Laboratory at KAUST. The role involves overseeing a high-performance computing cluster consisting of roughly 600 CPU and GPU nodes, managing storage systems, and handling InfiniBand and Ethernet networks. This position supports daily operations and assists a wide range of researchers in computational science, engineering, big data analytics, and AI/ML workloads.
Key Duties
- Deliver prompt and effective user assistance through various channels including phone, email, walk-in, and ticketing systems while upholding excellent customer service standards.
- Install, configure, and maintain HPC subsystems such as compute nodes, storage solutions, and networks using configuration management tools like Ansible or Puppet.
- Deploy and sustain cluster management and monitoring applications essential for HPC operations.
- Manage the Slurm workload manager by administering QOS policies, user accounts, accounting, and scripting automation in Python and C++.
- Create and update automation scripts primarily in Bash and Python to enhance system administration efficiency.
- Implement and support container platforms (such as Singularity/Apptainer and Docker) for HPC task execution.
- Conduct periodic performance benchmarking of CPUs, memory, network fabrics, and storage to optimize and tune hardware and software components.
- Apply security protocols including node hardening, kernel updates, and compliance maintenance across systems.
- Administer parallel file systems like Lustre, GPFS, Weka, or Vast, focusing on performance tuning and capacity management.
- Collaborate directly with academic and industrial partners to support computational, engineering, data analysis, and AI/ML research activities alongside application support teams.
- Develop internal software tools and utilities to facilitate research on cluster infrastructure.
- Lead proof-of-concept initiatives, evaluate new technologies, and promote system innovations.
- Coordinate with external vendors and service providers to troubleshoot and resolve hardware/software issues efficiently.
- Create and maintain comprehensive user guides, operational procedures, and training content in the internal documentation repository.
- Maintain up-to-date knowledge of HPC advancements, participate in relevant conferences, and spearhead benchmarking efforts to guide future infrastructure investments.
Required Competencies
- Extensive experience supporting users in computational science, engineering, data analytics, and AI applications in HPC contexts.
- Thorough knowledge of Linux system administration (RHEL, Rocky Linux, CentOS) within large-scale HPC environments.
- Proficient in HPC programming languages and models including Fortran, C/C++, Python, MPI, OpenMP, CUDA, and OpenACC.
- Proven background managing intricate HPC infrastructure including parallel file systems, job schedulers, high-speed networks, and monitoring tools.
- Hands-on experience with automation and configuration management tools such as Ansible or Puppet.
- Understanding of project management methodologies and proven ability to support research activities collaboratively.
- Strong analytical thinking, problem-solving aptitude, and effective decision-making capabilities.
- Self-driven with initiative to identify improvements and follow through to completion.
- Ability to manage multiple projects simultaneously, meeting deadlines with quality deliverables.
- Effective collaboration skills across diverse teams including researchers, application teams, and vendors within international multicultural settings.
- Excellent verbal and written English communication skills for technical reporting and presentations.
Qualifications and Experience
- Minimum education requirement: Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field.
- At least five years of experience in large-scale computing platform support and subsystem administration.
- Proficient in troubleshooting complex hardware problems and documenting detailed root cause analyses.
- Experienced managing parallel storage systems such as Lustre, GPFS, Weka, Vast, or equivalents.
- Skilled in performance benchmarking of CPU, memory, network, and storage components in HPC setups.
- Experienced administering workload managers like Slurm, LSF, or PBS.
- Strong Linux system administration skills specific to RHEL, Rocky Linux, or CentOS environments.
- Familiar with configuration management practices using tools like Ansible or Puppet.
- Demonstrated capacity to collaborate with researchers, support teams, and vendors to resolve difficult issues.
- Keen ability to drive initiatives to successful execution with cross-functional collaboration.
- Knowledge of Kubernetes and container orchestration platforms is advantageous.
Minimum education
Master's Degree