Technical Lead - GPU Infrastructure
Jobgether
وصف الوظيفة
About the role
This is a hands‑on technical leadership role responsible for architecting and delivering a large‑scale GPU infrastructure platform. You will lead the transition from managed Kubernetes workloads to bare‑metal GPU infrastructure, supporting research, model‑training and managed inference workloads.
Key responsibilities
- Own end‑to‑end platform architecture, including proposals, designs, technical reviews and documentation.
- Lead and line‑manage a distributed engineering team across backend, frontend, DevOps, QA and documentation.
- Design, build and operate a managed Slurm service for research and model‑training workloads.
- Lead GPU infrastructure operations on bare‑metal, handling NVIDIA drivers, CUDA, Fabric Manager, NVSwitch, DCGM and MIG.
- Bootstrap and maintain Kubernetes clusters on bare metal using NVIDIA GPU Operator and Network Operator.
- Define managed inference architecture covering serving, multi‑GPU parallelism, autoscaling and confidential‑compute capabilities.
- Establish observability across control plane, GPU fleet and applications with metrics, logging, alerting and SLOs.
- Lead incident response, on‑call model design and post‑incident reviews.
Required profile
- 8+ years of hands‑on engineering experience, with at least 3 years leading infrastructure platform teams.
- Bachelor’s or Master’s degree in Computer Science, Engineering or related field, or equivalent practical experience.
- Extensive production experience operating Slurm, HPC or GPU training clusters and NVIDIA GPU fleets on bare metal.
Required skills
- Slurm (slurmctld, slurmdbd, partitions, QoS, accounting, upgrades).
- Kubernetes (control plane, operators, multi‑tenancy, GPU Operator).
- NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG.
- InfiniBand, subnet configuration, RDMA, SR‑IOV, NCCL performance tuning.
- Linux kernel modules, PCIe passthrough, vfio‑pci, cgroups, namespaces.
- Observability tools: Prometheus, Grafana, Loki.
- JavaScript and Node.js for reviewing control‑plane and worker services.
- Experience with HPC storage (Lustre, VAST, NFS) and NVMe caching.
What we offer
- 100 % remote work with flexibility across UTC to UTC+5:30 time zones.
- Leadership of a sophisticated GPU infrastructure platform spanning bare‑metal, Slurm, Kubernetes and inference.
- Opportunity to shape architecture for advanced AI research, model training and managed inference workloads.
Questions fréquentes
لماذا تبلغ عن هذا العرض؟
اكتشف المزيد
الرواتب والأدلة وعمليات البحث في الإمارات العربية المتحدة.
الرواتب حسب المهنة
قدم طلبك في 30 ثانية
أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.
بالمتابعة، أنت توافق على شروط الاستخدام.
لديك حساب بالفعل؟ تسجيل الدخول
عزز فرصك
حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.
جاري تحليل سيرتك الذاتية...
Jobgether
عروض عمل ذات صلة
-
Web Manager
GEMS Education Dubai -
Technical Product Engineer – Backend / Full Stack (AI-Enabled)
Matajer AI Ajman -
Senior Software Engineer
Halian | Managed Services, Recruitment and Contract Staffing Abou Dabi -
QA Engineer - Automation Testing And Risk Management
Forsyth Barnes Consultancy Dubai -
Principal Engineer - Embedded Software
EDGE Abu Dhabi