📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Technical Lead - GPU Infrastructure

Jobgether

New Remote
Remote Senior 🇬🇧 English
Slurm Kubernetes NVIDIA drivers CUDA Fabric Manager NVSwitch DCGM MIG InfiniBand RDMA SR-IOV PCIe passthrough vfio-pci cgroups namespaces Prometheus Grafana Loki JavaScript Node.js Lustre NVMe caching

Job description

About the role

This is a hands‑on technical leadership role responsible for architecting and delivering a large‑scale GPU infrastructure platform. You will lead the transition from managed Kubernetes workloads to bare‑metal GPU infrastructure, supporting research, model‑training and managed inference workloads.

Key responsibilities

  • Own end‑to‑end platform architecture, including proposals, designs, technical reviews and documentation.
  • Lead and line‑manage a distributed engineering team across backend, frontend, DevOps, QA and documentation.
  • Design, build and operate a managed Slurm service for research and model‑training workloads.
  • Lead GPU infrastructure operations on bare‑metal, handling NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM and MIG.
  • Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
  • Define managed inference architecture covering serving, multi‑GPU parallelism, autoscaling and confidential‑compute capabilities.
  • Establish observability across control plane, GPU fleet and applications with metrics, logging, alerting and SLOs.
  • Lead incident response, on‑call model design and post‑incident reviews.

Required profile

  • 8+ years of hands‑on engineering experience, with at least 3 years leading infrastructure platform teams.
  • Bachelor’s or Master’s degree in Computer Science, Engineering or related field, or equivalent practical experience.
  • Extensive production experience operating Slurm, HPC or GPU training clusters and NVIDIA GPU fleets on bare metal.

Required skills

  • Slurm (slurmctld, slurmdbd, partitions, QoS, accounting, upgrades).
  • Kubernetes (control plane, operators, multi‑tenancy, GPU Operator).
  • NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG.
  • InfiniBand, subnet configuration, RDMA, SR‑IOV, NCCL performance tuning.
  • Linux kernel modules, PCIe passthrough, vfio‑pci, cgroups, namespaces.
  • Observability tools: Prometheus, Grafana, Loki.
  • JavaScript and Node.js for reviewing control‑plane and worker services.
  • Experience with HPC storage (Lustre, VAST, NFS) and NVMe caching.

What we offer

  • 100 % remote work with flexibility across UTC to UTC+5:30 time zones.
  • Leadership of a sophisticated GPU infrastructure platform spanning bare‑metal, Slurm, Kubernetes and inference.
  • Opportunity to shape architecture for advanced AI research, model training and managed inference workloads.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in the United Arab Emirates.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 13 hours ago

Expires 1 month from now

3 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Jobgether