📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Staff Engineer (Core & MLOps)

Jobgether

New Remote
Remote Senior 🇬🇧 English
Java Vert.x Netty Python gRPC Protocol Buffers Kubernetes Terraform Kafka Confluent Kafka Helm HAProxy Nginx Valkey Temporal DBOS MLOps SPIRE mTLS Cilium Istio Envoy

Job description

About the role

This role offers the opportunity to shape foundational infrastructure powering large‑scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI‑driven workflows to operate reliably and efficiently.

Key responsibilities

  • Architect and evolve control and context planes, including service registries, SLO enforcement, health‑aware routing, and automated canary releases.
  • Own the service chassis and golden path, maintaining multi‑language Java and Python client libraries, Helm charts, and deployment pipelines.
  • Define and govern inter‑service contracts such as gRPC and Protocol Buffer definitions, API gateway transcoding, and versioning policies.
  • Operate and improve core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, multi‑cluster routing, and automated failover.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and weighted canary deployments.

Required profile

  • 10+ years of experience building scalable distributed backend systems and internal platforms.
  • Advanced Java expertise with reactive frameworks (Vert.x or Netty) and strong Python proficiency.
  • Deep experience with gRPC, Protocol Buffers, and schema evolution in mission‑critical systems.
  • Hands‑on production experience with Kubernetes at scale, Terraform, and event‑streaming platforms such as Kafka.
  • Strong reliability engineering background, including SLO/SLI definition and fault tolerance.

Required skills

  • Java (Vert.x, Netty)
  • Python
  • gRPC & Protocol Buffers
  • Kubernetes
  • Terraform
  • Kafka (Confluent)
  • Helm
  • HAProxy / Nginx
  • Valkey
  • Temporal or DBOS (optional)
  • MLOps (model serving, performance monitoring)
  • Zero‑trust networking and service meshes (SPIRE, mTLS, Cilium, Istio, Envoy)

What we offer

  • Fully remote, remote‑first working environment with flexible hours.
  • Freedom to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large‑scale web data pipelines.
  • Exposure to cutting‑edge open‑source technologies and evolving AI infrastructure.
  • Opportunities to attend conferences and collaborate with a diverse, global engineering community.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in the United Arab Emirates.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 7 hours ago

Expires 1 month from now

6 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Jobgether