Technical Lead - GPU Infrastructure
Jobgether
Job description
About the role
This is a hands‑on technical leadership role responsible for architecting and delivering a large‑scale GPU infrastructure platform. You will lead the transition from managed Kubernetes workloads to bare‑metal GPU infrastructure, supporting research, model‑training and managed inference workloads.
Key responsibilities
- Own end‑to‑end platform architecture, including proposals, designs, technical reviews and documentation.
- Lead and line‑manage a distributed engineering team across backend, frontend, DevOps, QA and documentation.
- Design, build and operate a managed Slurm service for research and model‑training workloads.
- Lead GPU infrastructure operations on bare‑metal, handling NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM and MIG.
- Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
- Define managed inference architecture covering serving, multi‑GPU parallelism, autoscaling and confidential‑compute capabilities.
- Establish observability across control plane, GPU fleet and applications with metrics, logging, alerting and SLOs.
- Lead incident response, on‑call model design and post‑incident reviews.
Required profile
- 8+ years of hands‑on engineering experience, with at least 3 years leading infrastructure platform teams.
- Bachelor’s or Master’s degree in Computer Science, Engineering or related field, or equivalent practical experience.
- Extensive production experience operating Slurm, HPC or GPU training clusters and NVIDIA GPU fleets on bare metal.
Required skills
- Slurm (slurmctld, slurmdbd, partitions, QoS, accounting, upgrades).
- Kubernetes (control plane, operators, multi‑tenancy, GPU Operator).
- NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG.
- InfiniBand, subnet configuration, RDMA, SR‑IOV, NCCL performance tuning.
- Linux kernel modules, PCIe passthrough, vfio‑pci, cgroups, namespaces.
- Observability tools: Prometheus, Grafana, Loki.
- JavaScript and Node.js for reviewing control‑plane and worker services.
- Experience with HPC storage (Lustre, VAST, NFS) and NVMe caching.
What we offer
- 100 % remote work with flexibility across UTC to UTC+5:30 time zones.
- Leadership of a sophisticated GPU infrastructure platform spanning bare‑metal, Slurm, Kubernetes and inference.
- Opportunity to shape architecture for advanced AI research, model training and managed inference workloads.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Arab Emirates.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 13 hours ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Jobgether