Alle Stellen

NVIDIA is seeking a Senior GPU Networking Architect to join the networking software group, focusing on building and improving GPU communication kernels that connect GPU computing with networking. This role is integral to developing the software foundation for the largest AI systems globally, ensuring communication primitives are developed alongside GPU hardware capabilities.

Tasks

  • Build, implement, and optimize GPU communication kernels for collective and point-to-point operations in large-scale AI systems.
  • Leverage deep knowledge of GPU architecture, including thread scheduling, memory hierarchy, and execution pipelines, to improve kernel efficiency, minimize latency, and overlap computation with communication.
  • Develop GPU-resident communication primitives and device-side APIs for fine-grained, kernel-initiated data movement across nodes and accelerators.
  • Profile and tune GPU kernels end-to-end, identifying bottlenecks at the intersection of compute, memory, and network, and drive targeted optimizations.
  • Collaborate with network software, hardware, and AI framework teams to co-design communication strategies that align with GPU execution patterns and emerging model architectures.
  • Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and validate new communication strategies before production deployment.
  • Contribute to the evolution of programming models that expose GPU-aware networking capabilities to application developers.

Requirements

  • 5+ years of hands-on CUDA programming, including writing and optimizing non-trivial GPU kernels.
  • M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field.
  • Strong understanding of GPU architecture fundamentals: warp scheduling, shared memory, L2 cache, memory coalescing, occupancy tuning, and asynchronous execution.
  • Experience with systems-level C/C++ development in performance-critical environments.
  • Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication.
  • Ability to read and reason about GPU performance profiles (e.g., Nsight Compute, Nsight Systems) and translate observations into actionable optimizations.
  • Strong collaboration skills in a multi-national, interdisciplinary environment.
  • Experience developing or optimizing communication kernels in libraries such as NCCL, NVSHMEM, or similar GPU-aware communication frameworks.
  • Understanding of distributed deep learning parallelism techniques, including data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, and mixture-of-experts parallelism, and the communication patterns they impose on GPU kernels.
  • Background in RDMA, InfiniBand, high-speed networking, and GPU system topology, including NVLink, NVSwitch, PCIe, and network fabrics, and their impact on communication kernel design.
  • Experience with overlap techniques such as kernel pipelining, persistent kernels, or cooperative groups to hide communication latency behind compute.
  • Proven experience evaluating and optimizing large-scale LLM training or inference workloads, including hands-on work with frameworks such as PyTorch, TensorRT-LLM, or vLLM, and familiarity with emerging serving architectures such as disaggregated serving.

Benefits

  • Highly competitive salaries.
  • Comprehensive benefits package.
  • Base salary determined by location, experience, and pay of employees in similar positions.
  • For Poland: Base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.
Bist du Teil dieses Unternehmens?

Dieses Unternehmensprofil wurde automatisch erstellt. Wenn du für NVIDIA Switzerland AG arbeitest, kannst du das Profil jetzt beanspruchen und verifizieren – kostenlos und in wenigen Minuten.

Verifizierte Profile erhalten ein Siegel und können ihre Seite, Stellen und Bewerbungen direkt verwalten.
Über uns
NVIDIA Switzerland AG ist die Schweizer Gesellschaft des internationalen Technologieunternehmens NVIDIA. Das Unternehmen ist in der Entwicklung, Vermarktung und Bereitstellung von Technologien, Produkten und Dienstleistungen in Bereichen wie künstliche Intelligenz, beschleunigtes Computing, Grafikprozessoren, Robotik, virtuelle Realität, Informatik und professionelle Visualisierung tätig. NVIDIA entwickelt weltweit Chips, Systeme, Software und Plattformen für Rechenzentren, Forschung, Gaming, Automobilindustrie, Industrieanwendungen, Gesundheitswesen und weitere technologieintensive Branchen. Die Schweizer Gesellschaft unterstützt die Aktivitäten von NVIDIA in der Schweiz und ist im Bereich IT-Dienstleistungen registriert.
Das Team

Join a diverse, supportive environment with engineers developing the software foundation for the largest AI systems globally, collaborating across network software, hardware, and AI framework teams.

Ähnliche Stellen