Alle Stellen

We are seeking a Senior Solutions Architect with deep experience in large-scale production AI inference. You will collaborate with top EMEA AI Natives, AI infrastructure providers, and enterprises deploying AI at scale. As a trusted technical leader, you will ensure NVIDIA's inference stack achieves best performance, efficiency, and reliability, addressing the industry's toughest AI inference challenges at the intersection of AI and high-performance computing.

Tasks

  • Guide EMEA AI Natives customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters.
  • Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workload among thousands of GPUs.
  • Improve inference efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large MoE deployments.
  • Collaborate with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) to accelerate customer success.
  • Animate the AI inference developer’s community across EMEA through technical workshops, hackathons, and reference architectures.

Requirements

  • MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience.
  • 5+ years of experience in Neural Networks inference optimization.
  • Solid understanding of transformers inference optimization: quantization, disaggregated inference, speculative decoding, continuous batching, KV cache optimization.
  • Practical experience in MoE inference at scale: expert parallelism, WideEP, all-to-all communication, routing overhead, and load balancing at scale.
  • Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level.
  • Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling.
  • Understanding of GPU memory hierarchies and high-speed interconnects (NVLink, InfiniBand, RDMA, UCX).
  • Contribution to advanced AI labs or large scale AI infrastructure providers performing inference on thousands of GPUs.
  • Published work or benchmarks in large-scale AI inference.

Benefits

  • Highly competitive salaries.
  • Comprehensive benefits package.
Bist du Teil dieses Unternehmens?

Dieses Unternehmensprofil wurde automatisch erstellt. Wenn du für NVIDIA Switzerland AG arbeitest, kannst du das Profil jetzt beanspruchen und verifizieren – kostenlos und in wenigen Minuten.

Verifizierte Profile erhalten ein Siegel und können ihre Seite, Stellen und Bewerbungen direkt verwalten.
Über uns
NVIDIA Switzerland AG ist die Schweizer Gesellschaft des internationalen Technologieunternehmens NVIDIA. Das Unternehmen ist in der Entwicklung, Vermarktung und Bereitstellung von Technologien, Produkten und Dienstleistungen in Bereichen wie künstliche Intelligenz, beschleunigtes Computing, Grafikprozessoren, Robotik, virtuelle Realität, Informatik und professionelle Visualisierung tätig. NVIDIA entwickelt weltweit Chips, Systeme, Software und Plattformen für Rechenzentren, Forschung, Gaming, Automobilindustrie, Industrieanwendungen, Gesundheitswesen und weitere technologieintensive Branchen. Die Schweizer Gesellschaft unterstützt die Aktivitäten von NVIDIA in der Schweiz und ist im Bereich IT-Dienstleistungen registriert.
Ähnliche Stellen