Senior Solutions Architect – Large Scale Neural Networks Inference
nvidia.wd5.myworkdayjobs.com
- Job written in
- English
- Location
- Switzerland
- Work type
- On-site
- Type
- Full-time
NVIDIA is looking for a Senior Solutions Architect who will shape large‑scale neural network inference across its EMEA customer base. The role sits within a company that pioneered accelerated computing and now powers AI infrastructure for a wide range of industries worldwide. In this position you will own the inference strategy for a portfolio of "AI Natives" customers, steering projects from early proof‑of‑concept through to full‑scale production. You will pinpoint performance bottlenecks such as latency, token cost, memory use and networking delays, then design and fine‑tune high‑throughput pipelines with tools like NVIDIA Dynamo, TensorRT‑LLM, vLLM and SGLang. Your work will translate real‑world deployment patterns into concrete feedback that informs the roadmap for NVIDIA's inference stack, including Dynamo, TensorRT‑LLM and NIM, while coordinating closely with both internal teams and client engineers. The essential qualifications include an MS or PhD in Computer Science, Engineering or a comparable background, and at least eight years of experience building AI/ML infrastructure. Candidates must possess deep expertise in optimizing LLM/VLM inference at scale, with a solid grasp of transformer acceleration techniques such as INT4/FP8 quantisation, speculative decoding, disaggregated inference, continuous batching, KV‑cache optimisation and WideEP for MoE models. A thorough understanding of GPU memory hierarchies and low‑latency networking, a proven record of leading technical initiatives, and strong communication skills for interacting with research scientists, infrastructure engineers and senior executives are also required. Preferred experience includes hands‑on work with NVIDIA's own inference ecosystem, TensorRT‑LLM, Triton Inference Server, NIM and Dynamo, as well as managing GPU workloads on Kubernetes. Candidates who have run inference at scale inside a frontier AI lab or a hyperscale inference team, or who have contributed to open‑source projects such as vLLM, SGLang, KServe or NVIDIA Dynamo, will stand out. The role is based in Poland, with a salary band of 292,500 – 507,000 PLN for Level 4 and 375,000 – 650,000 PLN for Level 5, complemented by NVIDIA's competitive benefits package and a strong commitment to diversity and equal opportunity. What the role asks for: - MS or PhD in Computer Science or Engineering - 8+ years AI/ML infrastructure experience - Expertise in LLM/VLM inference optimization at scale - Knowledge of transformer acceleration (INT4/FP8, speculative decoding, etc.) - Understanding of GPU memory hierarchies and low‑latency networking - Proven ability to lead technical initiatives - Excellent communication with scientists, engineers, executives - Experience with NVIDIA TensorRT‑LLM, Triton, NIM, Dynamo (nice‑to‑have) - GPU orchestration on Kubernetes (nice‑to‑have) - Inference work in frontier AI lab or hyperscale team (nice‑to‑have) - Contributions to open‑source projects vLLM, SGLang, KServe, Dynamo (nice‑to‑have)
Lebenslauf-Vorlage für deine Bewerbung
Für diesen Titel gibt es keine eigene Vorlage — aber über 190 nach Beruf, alle im Schweizer Aufbau mit Beispieltext. Nimm die, die deiner Stelle am nächsten kommt.
Lebenslauf-Vorlage Schweiz ansehenWeitersuchen
- Alle Architekt Jobs in der Schweiz
- Lebenslauf mit KI erstellen
- ATS-Lebenslauf prüfen
- Bewerbungsschreiben für dieses Inserat