Staff Engineer (Core & MLOps)
Jobgether
Job description
About the role
This role offers the opportunity to shape foundational infrastructure powering large‑scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI‑driven workflows to operate reliably and efficiently.
Key responsibilities
- Architect and evolve control and context planes, including service registries, SLO enforcement, health‑aware routing, and automated canary releases.
- Own the service chassis and golden path, maintaining multi‑language Java and Python client libraries, Helm charts, and deployment pipelines.
- Define and govern inter‑service contracts such as gRPC and Protocol Buffer definitions, API gateway transcoding, and versioning policies.
- Operate and improve core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, Valkey, and database modernization initiatives.
- Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, multi‑cluster routing, and automated failover.
- Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and weighted canary deployments.
Required profile
- 10+ years of experience building scalable distributed backend systems and internal platforms.
- Advanced Java expertise with reactive frameworks (Vert.x or Netty) and strong Python proficiency.
- Deep experience with gRPC, Protocol Buffers, and schema evolution in mission‑critical systems.
- Hands‑on production experience with Kubernetes at scale, Terraform, and event‑streaming platforms such as Kafka.
- Strong reliability engineering background, including SLO/SLI definition and fault tolerance.
Required skills
- Java (Vert.x, Netty)
- Python
- gRPC & Protocol Buffers
- Kubernetes
- Terraform
- Kafka (Confluent)
- Helm
- HAProxy / Nginx
- Valkey
- Temporal or DBOS (optional)
- MLOps (model serving, performance monitoring)
- Zero‑trust networking and service meshes (SPIRE, mTLS, Cilium, Istio, Envoy)
What we offer
- Fully remote, remote‑first working environment with flexible hours.
- Freedom to work from the location where you are most productive.
- Opportunity to work on core infrastructure supporting large‑scale web data pipelines.
- Exposure to cutting‑edge open‑source technologies and evolving AI infrastructure.
- Opportunities to attend conferences and collaborate with a diverse, global engineering community.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Arab Emirates.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 7 hours ago
Expires 1 month from now
6 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Jobgether
Related job offers
-
Software Engineer (Orchestration)
Callsign Abou Dabi -
Software Engineer (Full Stack)
Callsign Abou Dabi -
IT Administrator – Hotel Technology Support
St. Regis Hotels & Resorts Abou Dabi -
Quality Lead – Enterprise Applications
Indixpert Doubaï -
Service Desk Supervisor / Team Lead (Arabic Required) – Dubai
Sanvi Engineering And Business Consulting Solutions Doubaï