G

Orchestration Workload Engineer - ACE - AI Factory

Genentech·Kaiseraugst·18.09.2026

 careers.gene.com

 
 
full_time80–100%
Required language
English conversational
Job written in
English
Location
Kaiseraugst
Work type
Hybrid
Type
Full-time

 

The position of Workload Orchestration Engineer sits within Roche's Accelerated Compute Engineering (ACE) team. ACE acts as a centre of excellence for high‑performance compute and AI infrastructure, supporting Roche sites worldwide with on‑premises and cloud‑based platforms that enable research, development and digital product delivery. In day‑to‑day work you will be the recognised authority on the SLURM Workload Manager, responsible for designing, scaling and maintaining SLURM across heterogeneous CPU and GPU clusters. You will tune advanced configurations, integrate custom plugins, manage topology‑aware scheduling and implement complex QoS and fair‑share policies. The role also requires bridging SLURM with Kubernetes‑based solutions, applying container runtimes such as Singularity or Apptainer, and resolving multi‑tenant bottlenecks that affect GPU allocation, MPI/NCCL communication and job reliability. Collaboration with observability engineers to build telemetry dashboards and with global cross‑functional teams to define orchestration standards is a core part of the job. The essential qualifications include a bachelor's or higher degree in computer science, applied mathematics, computational engineering or a closely related field. You must have extensive systems‑engineering experience focused on workload scheduling, SLURM administration and optimisation of multi‑tenant clusters. A proven record of leading complex technical projects and mentoring junior engineers is required, as is prior experience working in life‑science, pharmaceutical R&D or high‑performance scientific research environments. Additional strengths that are valued comprise deep expertise in architecting and upgrading production SLURM installations, strong knowledge of SlurmDBD accounting and scheduler telemetry, and practical experience with Kubernetes and container runtimes (Singularity, Apptainer, Docker, Enroot). Familiarity with GPU scheduling technologies (NVIDIA MIG), high‑speed interconnects (InfiniBand, RoCE) and communication frameworks (MPI, NCCL) is advantageous, as is advanced proficiency with Infrastructure‑as‑Code tools such as Ansible or Terraform. Broad understanding of HPC, AI infrastructure, networking and observability further supports success in this globally distributed team. What the role asks for: - Bachelor's or advanced degree in CS, Applied Math, Computational Engineering or related - Extensive systems engineering experience in workload scheduling and SLURM administration - Proven track record leading complex technical initiatives and mentoring peers - Experience in life sciences, pharmaceutical R&D or high‑performance scientific research - Expertise in architecting, scaling and optimizing production SLURM environments (nice‑to‑have) - Deep knowledge of SlurmDBD, accounting and scheduler telemetry (nice‑to‑have) - Hands‑on experience with Kubernetes and container runtimes (Singularity, Apptainer, Docker) (nice‑to‑have) - Familiarity with GPU scheduling, NVIDIA MIG and high‑speed interconnects (nice‑to‑have) - Proficiency with Infrastructure‑as‑Code tools such as Ansible or Terraform (nice‑to‑have) - Broad understanding of HPC, AI infrastructure, networking and observability (nice‑to‑have)

 

 

 

 

 

Kostenlos

Engineer: Lebenslauf-Vorlage

Bewirb dich mit einem Lebenslauf im Schweizer Aufbau — mit Beispieltext und den Anforderungen, die Engineer-Inserate am häufigsten nennen.

Lebenslauf-Vorlage Schweiz ansehen

Engineer: was der Markt gerade verlangt

Engineer-Stellen gehören zu den regelmässig ausgeschriebenen Berufen auf SwissJobs.app.

Weitersuchen

Ähnliche Jobs per E-Mail erhalten