📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Staff Site Reliability Engineer

Jobgether

New Remote
Remote Senior 🇬🇧 English
Observability Incident Management Reliability Metrics Automation

Job description

About the role

We are seeking a Staff Site Reliability Engineer to lead reliability initiatives across multiple engineering teams in a fully remote, globally distributed organization. As the first dedicated SRE, you will shape reliability practices, embed SRE principles into the engineering culture, and drive measurable improvements in system resilience.

Key responsibilities

  • Define and implement SLIs, SLOs, and error budgets for critical production paths.
  • Establish reliability metrics and dashboards for engineering leadership.
  • Strengthen the incident management lifecycle, including detection, response, communication, post‑mortems, and follow‑up actions.
  • Improve alert quality, anomaly detection, and escalation processes in partnership with infrastructure teams.
  • Lead reliability assessments for high‑risk changes, covering production readiness, capacity, failure modes, and rollback strategies.
  • Introduce deliberate failure testing, game days, and chaos engineering exercises.
  • Coach engineers to become reliability advocates and promote distributed SRE ownership.
  • Develop lightweight operational standards for on‑call practices, runbooks, and change safety.
  • Collaborate with architects to embed reliability and failure tolerance into system design.
  • Remain hands‑on during incidents, building tooling, dashboards, and automation as needed.

Required profile

  • Proven senior‑level experience in Site Reliability Engineering or related fields.
  • Strong leadership and influence across multiple engineering teams.
  • Hands‑on mindset with the ability to build tooling and automation.
  • Excellent communication skills for incident coordination and post‑mortem facilitation.

Required skills

  • Site Reliability Engineering (SRE) practices
  • Observability and monitoring
  • Incident management and response
  • Reliability metrics (SLI/SLO/error budgets)
  • Chaos engineering and failure testing
  • Automation and tooling development
  • Capacity planning and production readiness

What we offer

  • Full remote work environment with global collaboration.
  • High autonomy to define standards and shape reliability culture.
  • Opportunity to lead SRE initiatives at a senior level.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Jobgether.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 21 hours ago

Expires 1 month from now

9 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Jobgether