Senior Site Reliability Engineer, DGX Cloud
nvidia.wd5.myworkdayjobs.com
- Job written in
- English
- Location
- Zürich
- Work type
- On-site
- Type
- Full-time
NVIDIA, a pioneer in graphics, gaming and accelerated computing for more than a quarter‑century, is expanding its AI platform through DGX Cloud, a fully managed service that runs on the major public clouds and on‑premise infrastructure. As a Senior Site Reliability Engineer on the DGX Cloud team you will help keep the high‑performance clusters that serve AI researchers and enterprise customers reliable and performant worldwide. Your daily work will involve designing, building and supporting large‑scale Kubernetes environments, with a strong emphasis on performance, real‑time monitoring, logging and alerting. You will define service level objectives and indicators, track error budgets, and produce regular health reports. Before new services go live you will assist with system design, create tooling, manage capacity and conduct launch reviews, then continue to monitor availability, latency and overall health once they are in production. The role also includes operating and tuning GPU workloads across AWS, GCP, Azure, OCI and private clouds, leading incident triage, performing root‑cause analysis, conducting blameless post‑mortems and participating in an on‑call rotation. Candidates must hold a bachelor’s degree in computer science or a related field, or possess equivalent experience, and bring at least ten years of hands‑on production service operation. Required expertise includes deep Kubernetes administration, containerization and micro‑services, infrastructure automation (such as Terraform or Ansible), a high‑level language like Python or Go, Linux systems, networking fundamentals and cloud security standards. A solid grasp of SRE practices—SLOs, SLIs, error budgets and incident handling—and experience building observability stacks with tools such as Prometheus, Grafana, OpenTelemetry or the ELK suite are essential. Experience that distinguishes applicants includes running GPU‑accelerated clusters with KubeVirt, applying generative‑AI methods to reduce operational toil, or working with workflow orchestration platforms like Temporal, Airflow or Argo. The position offers the chance to influence a global AI infrastructure while collaborating with a diverse, supportive team at the forefront of high‑performance computing.
Engineer: Lebenslauf-Vorlage
Bewirb dich mit einem Lebenslauf im Schweizer Aufbau — mit Beispieltext und den Anforderungen, die Engineer-Inserate am häufigsten nennen.
Lebenslauf-Vorlage Schweiz ansehenEngineer: was der Markt gerade verlangt
Engineer-Stellen gehören zu den regelmässig ausgeschriebenen Berufen auf SwissJobs.app.
Weitersuchen
- Alle DevOps Engineer Jobs in der Schweiz
- Alle Jobs in Zürich
- Lebenslauf als Engineer: Beispiel und Vorlage
- Lebenslauf mit KI erstellen
- ATS-Lebenslauf prüfen
- Bewerbungsschreiben für dieses Inserat