Jobiglo

لا توجد نتائج.

Technical Lead - GPU Infrastructure

Tether.io

جديد Remote
Remote Senior 🇬🇧 English
Slurm Kubernetes Linux NVIDIA driver CUDA InfiniBand RDMA Prometheus Grafana Loki JavaScript Node.js HPC storage DCGM MIG

وصف الوظيفة

About the role

Join Tether Data's Cosmic AC GPU compute and managed inference platform as Technical Lead. You will own the architecture and delivery of a full‑stack bare‑metal GPU infrastructure, leading a distributed team of about twelve engineers across backend, frontend, DevOps, QA and documentation.

Key responsibilities

  • Own the end‑to‑end platform architecture, producing high‑level and low‑level designs and keeping them up to date.
  • Lead and line‑manage a distributed team, establishing engineering standards, conducting code reviews, and driving growth.
  • Design, build and operate a managed Slurm scheduling layer for research and model‑training workloads.
  • Bootstrap and maintain the Kubernetes control plane on partner‑provided bare‑metal, handling GPU operator, network operator, upgrades and day‑2 operations.
  • Develop and scale managed inference services, including multi‑GPU, multi‑node parallelism and confidential‑compute capabilities.
  • Implement observability across control plane, GPU fleet and applications using metrics, logging and SLOs, and run an on‑call rotation.
  • Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests.
  • Collaborate with internal research, model‑training and product teams to align platform capabilities with workload needs.
  • Recruit and set technical standards for new engineers joining the platform team.

Required profile

  • 8+ years of hands‑on engineering experience, with at least 3 years leading teams that build and operate critical infrastructure platforms.
  • Bachelor’s or Master’s degree in Computer Science, Engineering or equivalent practical experience.
  • Proven experience running Slurm at scale for real users, including partitions, QoS, accounting and node‑health automation.
  • Deep knowledge of GPU fleet operation on bare metal: NVIDIA driver, CUDA lifecycle, MIG, DCGM health monitoring and node burn‑in.
  • Expertise in high‑performance interconnects such as InfiniBand, RDMA and SR‑IOV, and ability to troubleshoot multi‑node NCCL performance.
  • Strong Linux systems background, including kernel modules, PCIe passthrough, cgroups, namespaces and performance tuning.
  • Production‑grade Kubernetes operations experience, covering control‑plane upgrades, CNI/CSI, operators and multi‑tenancy design.
  • Hands‑on experience with HPC storage solutions (Lustre, NFS, VAST) and large‑scale data movement.
  • Fluency in JavaScript/Node.js sufficient to review and influence control‑plane services.

Required skills

  • Slurm
  • Kubernetes
  • Linux system administration
  • NVIDIA driver & CUDA
  • InfiniBand, RDMA, SR‑IOV
  • Prometheus, Grafana, Loki (observability)
  • JavaScript, Node.js
  • HPC storage (Lustre, NFS, VAST)
  • GPU fleet management (DCGM, MIG)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Tether.io.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

لماذا تبلغ عن هذا العرض؟

شكراً لإبلاغك. سنراجع هذا العرض.

اكتشف المزيد

الرواتب والأدلة وعمليات البحث في مصر.

قدم طلبك في 30 ثانية

أدخل بريدك الإلكتروني للتقديم. سيتم إنشاء حساب تلقائياً.

بالمتابعة، أنت توافق على شروط الاستخدام.

لديك حساب بالفعل؟ تسجيل الدخول

💬 راسلنا على تيليجرام الدردشة عبر واتساب

منشور منذ 10 ساعات

ينتهي شهر من الآن

10 مشاهدات · 0 مهتم

عزز فرصك

حمّل سيرتك الذاتية وسنقترح عليك الوظائف التي تناسب ملفك.

جاري تحليل سيرتك الذاتية...

Tether.io