Jobiglo

لا توجد نتائج.

Technical Lead - GPU Infrastructure

Jobgether

جديد Remote
Remote Senior 🇬🇧 English
Slurm Kubernetes NVIDIA drivers CUDA Fabric Manager NVSwitch DCGM MIG InfiniBand RDMA SR-IOV PCIe passthrough vfio-pci cgroups namespaces Prometheus Grafana Loki JavaScript Node.js Lustre NVMe caching

وصف الوظيفة

About the role

This is a hands‑on technical leadership role responsible for architecting and delivering a large‑scale GPU infrastructure platform. You will lead the transition from managed Kubernetes workloads to bare‑metal GPU infrastructure, supporting research, model‑training and managed inference workloads.

Key responsibilities

  • Own end‑to‑end platform architecture, including proposals, designs, technical reviews and documentation.
  • Lead and line‑manage a distributed engineering team across backend, frontend, DevOps, QA and documentation.
  • Design, build and operate a managed Slurm service for research and model‑training workloads.
  • Lead GPU infrastructure operations on bare‑metal, handling NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM and MIG.
  • Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
  • Define managed inference architecture covering serving, multi‑GPU parallelism, autoscaling and confidential‑compute capabilities.
  • Establish observability across control plane, GPU fleet and applications with metrics, logging, alerting and SLOs.
  • Lead incident response, on‑call model design and post‑incident reviews.

Required profile

  • 8+ years of hands‑on engineering experience, with at least 3 years leading infrastructure platform teams.
  • Bachelor’s or Master’s degree in Computer Science, Engineering or related field, or equivalent practical experience.
  • Extensive production experience operating Slurm, HPC or GPU training clusters and NVIDIA GPU fleets on bare metal.

Required skills

  • Slurm (slurmctld, slurmdbd, partitions, QoS, accounting, upgrades).
  • Kubernetes (control plane, operators, multi‑tenancy, GPU Operator).
  • NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG.
  • InfiniBand, subnet configuration, RDMA, SR‑IOV, NCCL performance tuning.
  • Linux kernel modules, PCIe passthrough, vfio‑pci, cgroups, namespaces.
  • Observability tools: Prometheus, Grafana, Loki.
  • JavaScript and Node.js for reviewing control‑plane and worker services.
  • Experience with HPC storage (Lustre, VAST, NFS) and NVMe caching.

What we offer

  • 100 % remote work with flexibility across UTC to UTC+5:30 time zones.
  • Leadership of a sophisticated GPU infrastructure platform spanning bare‑metal, Slurm, Kubernetes and inference.
  • Opportunity to shape architecture for advanced AI research, model training and managed inference workloads.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.

اكتشف المزيد

الرواتب والأدلة وعمليات البحث في الإمارات العربية المتحدة.

قدم طلبك في 30 ثانية

أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.

بالمتابعة، أنت توافق على شروط الاستخدام.

لديك حساب بالفعل؟ تسجيل الدخول

💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ يوم

ينتهي شهر من الآن

11 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Jobgether