Lead HPC Engineer – AI Infrastructure
Osmii · Abu Dhabi
Job description
About the role
We are currently partnered with a premier global hyperscaler that is rapidly expanding its AI digital infrastructure footprint in Abu Dhabi, UAE. We are looking for a Lead HPC Engineer to architect, optimise, and scale ultra‑low latency, custom hardware clusters powering next‑generation AI, machine learning, and scientific workloads.
Key responsibilities
- Architect & deploy large‑scale GPU/accelerator clusters for enterprise AI training.
- Optimise ultra‑low latency fabrics using RoCEv2, RDMA, and custom network topologies.
- Tune parallel file systems (Lustre, GPFS) and Linux kernel stacks for maximum throughput.
- Drive automation across Slurm, Kubernetes, Python, and C/C++ environments.
Required profile
- 5+ years in HPC, hyperscale infrastructure, or supercomputing environments.
- Deep expertise in low‑latency networking, parallel storage, and workload managers.
- Strong background in performance profiling, hardware optimisation, and automation.
Required skills
- GPU/accelerator clusters
- RoCEv2
- RDMA
- Custom network topologies
- Lustre
- GPFS
- Linux kernel
- Slurm
- Kubernetes
- Python
- C/C++
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Arab Emirates.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 10 hours ago
Expires 1 month from now
6 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Osmii
Abu Dhabi