Senior Solutions Architect – Large Scale AI Training
nvidia.wd5.myworkdayjobs.com
- Required language
- English conversational
- Job written in
- English
- Location
- Switzerland
- Work type
- On-site
- Type
- Full-time
The role of Senior Solutions Architect, Large Scale AI Training at NVIDIA involves supporting top AI organizations and research institutions in the EMEA region. The company is a leader in accelerated computing and AI infrastructure, powering global intelligence across various industries. The position focuses on large-scale distributed training and alignment of Neural Networks, particularly with large Mixture-of-Experts (MoE) models. The main day-to-day responsibilities include building and managing strategic technical relationships with leading AI model builders. This involves collaborating closely with customers to define the software stack and infrastructure required for large-scale training and post-training workflows, including Reinforcement Learning (RL). The role also entails serving as the go-to expert on distributed training strategies, guiding customers through efficient large-scale training and post-training recipes using Deep Learning Frameworks such as PyTorch, Megatron-LM, or NeMo (RL/Gym). Additionally, the role involves helping customers optimize training/finetuning efficiency at scale, covering GPU utilization, communication overlap, and memory management. The architect will also collaborate with NVIDIA's product and research teams to present customer needs and animate the developer community by building or supporting hackathons, demos, and technical conferences. To qualify for this role, candidates must have an MS or PhD in Computer Science, Engineering, or equivalent experience. Over 7 years of practical experience in distributed AI training, including direct involvement with HPC and/or AI environments with multi-node GPU clusters, is required. A solid understanding of training infrastructure and how it affects efficiency and scalability is essential. Strong proficiency with Megatron-LM, NeMo, or equivalent distributed training frameworks is also necessary. Excellent communication skills are crucial for engaging both research scientists and infrastructure engineers. Nice-to-have skills include experience in fine-tuning with Reinforcement Learning (RLVR, RLHF) at scale, experience with LatentMoE, expert load balancing, and speculative decoding for MoE inference. Prior experience in an AI Datacenter/HPC center, national lab, or frontier AI lab environment is also beneficial. Published work or open-source contributions in distributed training can further enhance a candidate's profile. The role is based in Poland, with a base salary range of 292,500 PLN - 507,000 PLN. What the role asks for: - MS or PhD in Computer Science, Engineering, or equivalent experience. - Over 7 years of practical experience in distributed AI training, including direct involvement with HPC and/or AI environments with multi-node GPU clusters. - Solid understanding of training infrastructure and how it affects efficiency and scalability. - Strong proficiency with Megatron-LM, NeMo, or equivalent distributed training frameworks. - Excellent communication skills with an ability to engage both research scientists and infrastructure engineers. - Experience in fine-tuning with Reinforcement Learning (RLVR, RLHF) at scale. (Nice to have) - Experience with LatentMoE, expert load balancing, and speculative decoding for MoE inference. (Nice to have) - Prior experience in an AI Datacenter/HPC center, national lab, or frontier AI lab environment. (Nice to have) - Published work or open-source contributions in distributed training. (Nice to have)
Lebenslauf-Vorlage für deine Bewerbung
Für diesen Titel gibt es keine eigene Vorlage — aber über 190 nach Beruf, alle im Schweizer Aufbau mit Beispieltext. Nimm die, die deiner Stelle am nächsten kommt.
Lebenslauf-Vorlage Schweiz ansehenWeitersuchen
- Alle Architekt Jobs in der Schweiz
- Lebenslauf mit KI erstellen
- ATS-Lebenslauf prüfen
- Bewerbungsschreiben für dieses Inserat