IT Administrator HPC & Data Center Operations
Prepaire Labs · Abu Dhabi
Job description
About the role
Prepaire Labs is looking for an experienced IT Administrator to oversee its High‑Performance Computing (HPC) and data‑center infrastructure that supports AI‑driven drug discovery and healthcare research. The role is based on‑site in Abu Dhabi and requires hands‑on management of CPU/GPU clusters, networking, storage and software licensing to ensure high availability, security and regulatory compliance.
Key responsibilities
- Administer and monitor HPC clusters with CPU and GPU nodes (e.g., NVIDIA H100/A100/L40S, AMD EPYC, Intel Xeon).
- Deploy and maintain workload managers and job schedulers such as SLURM, PBS or Kubernetes with GPU orchestration.
- Manage GPU drivers, CUDA toolkits, container runtimes (Docker, Singularity/Apptainer) and associated ML frameworks.
- Oversee data‑center operations including rack‑and‑stack, power, cooling, cable management, hardware lifecycle and capacity planning.
- Perform preventive maintenance, firmware/BIOS updates, hardware diagnostics and coordinate vendor support cases.
- Maintain high‑performance storage systems (NFS, Lustre, BeeGFS, Ceph) and backup/disaster‑recovery strategies.
- Monitor cluster health and performance using Grafana, Prometheus, Zabbix, Nagios or NVIDIA DCGM and optimise resource allocation.
- Design, configure and maintain LAN/WAN, VLANs, firewalls, VPNs and high‑speed interconnects (InfiniBand, RoCE, 10/25/40/100 GbE).
- Administer core networking equipment from Cisco, Juniper, Arista, Fortinet and Mellanox/NVIDIA.
Required profile
- Minimum 60 months (5 years) of experience in HPC and data‑center operations.
- Proven ability to work on‑site in a data‑center environment.
- Strong problem‑solving skills and experience with hardware lifecycle management.
Required skills
- GPU computing stacks
- Data Center Operations
- Software licensing management
- Linux Administration
- NVIDIA H100, A100, L40S
- AMD EPYC, Intel Xeon
- SLURM, PBS, Kubernetes
- Docker, Singularity/Apptainer, NVIDIA Container Toolkit
- Grafana, Prometheus, Zabbix, Nagios, NVIDIA DCGM
- InfiniBand, RoCE, 10GbE, 25GbE, 40GbE, 100GbE
- Cisco, Juniper, Arista, Fortinet, Mellanox, NVIDIA networking
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Arab Emirates.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 1 week from now
47 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Prepaire Labs
Abu Dhabi