Senior Solutions Architect – Large Scale AI Inference
nvidia.wd5.myworkdayjobs.com
- Required language
- English conversational
- Job written in
- English
- Location
- Switzerland
- Work type
- On-site
- Type
- Full-time
NVIDIA is seeking a Senior Solutions Architect focused on large‑scale AI inference. As a leading provider of accelerated computing, NVIDIA powers AI infrastructure across industries worldwide. The role sits within the EMEA region, working closely with AI‑native customers, infrastructure partners and enterprise teams that run demanding AI workloads at scale. In this position you will guide customers through the deployment and optimisation of inference workloads on multi‑node GPU clusters. You will design efficient pipelines for both dense and sparse/latent Mixture‑of‑Experts (MoE) models, handling thousands of GPUs. Key tasks include improving performance through INT4/FP8 quantisation, speculative decoding, disaggregated pre‑fill/decode, KV‑cache management and WideEP techniques. You will partner with NVIDIA product groups such as Dynamo, TensorRT‑LLM and NIXL to accelerate customer success, and you will nurture the AI inference developer community in EMEA by delivering workshops, hackathons and reference architectures. The ideal candidate holds an MS or PhD in Computer Science, Engineering, High‑Performance Computing or equivalent experience, and brings at least five years of hands‑on work in neural‑network inference optimisation. A solid grasp of transformer inference, including quantisation, disaggregated inference, speculative decoding, continuous batching and KV‑cache optimisation, is required. Practical experience with MoE inference at scale, covering expert parallelism, WideEP, all‑to‑all communication, routing overhead and load‑balancing, is essential. Strong communication skills are needed to engage deeply with ML engineers, researchers and systems architects. Additional strengths that set candidates apart include direct experience with NVIDIA's Dynamo, NIXL or Grove toolsets, and a deep understanding of GPU memory hierarchies and high‑speed interconnects such as NVLink, InfiniBand, RDMA and UCX. Contributions to advanced AI labs or large‑scale AI infrastructure providers, as well as published benchmarks or research on massive AI inference, are valued. The role is based in Poland, offering a salary range of 292,500 – 507,000 PLN, competitive benefits and a commitment to diversity and equal opportunity. What the role asks for: - MS or PhD in Computer Science, Engineering, or HPC - 5+ years neural network inference optimization experience - Deep knowledge of transformer inference quantization and optimization - Experience with MoE inference at scale (expert parallelism, WideEP) - Ability to engage with ML engineers, researchers, architects - Hands‑on experience with NVIDIA Dynamo, NIXL, or Grove (nice‑to‑have) - Understanding of GPU memory hierarchies and NVLink/InfiniBand (nice‑to‑have) - Published work or benchmarks in large‑scale AI inference (nice‑to‑have)
Lebenslauf-Vorlage für deine Bewerbung
Für diesen Titel gibt es keine eigene Vorlage — aber über 190 nach Beruf, alle im Schweizer Aufbau mit Beispieltext. Nimm die, die deiner Stelle am nächsten kommt.
Lebenslauf-Vorlage Schweiz ansehenWeitersuchen
- Alle Architekt Jobs in der Schweiz
- Lebenslauf mit KI erstellen
- ATS-Lebenslauf prüfen
- Bewerbungsschreiben für dieses Inserat