📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Senior LLMOps / AI Platform Engineer

Hire Rightt - Executive Search & HR Advisory · Doubaï

New
Senior 25,000 - 30,000 AED/month 🇬🇧 English
Python FastAPI vLLM Hugging Face Kubernetes Docker Helm AWS NVIDIA GPU inference performance optimization LangChain LangGraph LangSmith Langfuse RAG embeddings Qdrant Milvus OpenTelemetry Prometheus Grafana PostgreSQL Redis GitHub Actions CI/CD

Job description

About the role

The Senior LLMOps / AI Platform Engineer will design, build, and operate production‑grade large language model (LLM) and generative AI infrastructure for a financial services firm.

Key responsibilities

  • Deploy and operate self‑hosted LLMs using vLLM, SGLang and Ollama.
  • Optimize LLM inference for latency, throughput, concurrency, GPU memory, KV cache and cost.
  • Manage GPU workloads across multiple NVIDIA GPUs.
  • Deploy AI services on Kubernetes / AWS EKS with Docker and Helm.
  • Implement reliability mechanisms, health checks, monitoring, automated recovery and model refresh strategies.
  • Set up observability using Langfuse/LangSmith, OpenTelemetry, Prometheus and Grafana.
  • Deploy and optimise Retrieval‑Augmented Generation (RAG) systems, embedding models and vector databases such as Qdrant and Milvus.
  • Support AI agents and workflows built with LangChain and LangGraph.
  • Build and maintain CI/CD pipelines for AI services and infrastructure.
  • Troubleshoot production issues across LLMs, GPUs, Kubernetes, networking and AI applications.

Required profile

  • Strong experience with Python and FastAPI.
  • Hands‑on expertise in deploying self‑hosted LLMs and using Hugging Face models.
  • Deep knowledge of Kubernetes, Docker, Helm and AWS cloud services.
  • Proven ability to optimise NVIDIA GPU inference performance.
  • Experience with LangChain, LangGraph and LangSmith tooling.

Required skills

  • Python, FastAPI
  • vLLM, Hugging Face, self‑hosted LLM deployment
  • Kubernetes, Docker, Helm, AWS
  • NVIDIA GPU inference and performance tuning
  • LangChain, LangGraph, LangSmith, Langfuse
  • RAG, embeddings, vector databases (Qdrant, Milvus)
  • Observability: OpenTelemetry, Prometheus, Grafana
  • PostgreSQL, Redis
  • GitHub Actions, CI/CD pipelines

Questions fréquentes

Le salaire proposé pour ce poste est de 25-30k AED par mois. Le détail figure dans l'annonce.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 4 hours ago

Expires 1 month from now

6 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Hire Rightt - Executive Search & HR Advisory

Doubaï