Jobiglo

No results.

Technical Lead - GPU Infrastructure

Tether.io

New Remote
Remote Senior 🇬🇧 English
Slurm Kubernetes Linux NVIDIA driver CUDA InfiniBand RDMA Prometheus Grafana Loki JavaScript Node.js HPC storage DCGM MIG

Job description

About the role

Join Tether Data's Cosmic AC GPU compute and managed inference platform as Technical Lead. You will own the architecture and delivery of a full‑stack bare‑metal GPU infrastructure, leading a distributed team of about twelve engineers across backend, frontend, DevOps, QA and documentation.

Key responsibilities

  • Own the end‑to‑end platform architecture, producing high‑level and low‑level designs and keeping them up to date.
  • Lead and line‑manage a distributed team, establishing engineering standards, conducting code reviews, and driving growth.
  • Design, build and operate a managed Slurm scheduling layer for research and model‑training workloads.
  • Bootstrap and maintain the Kubernetes control plane on partner‑provided bare‑metal, handling GPU operator, network operator, upgrades and day‑2 operations.
  • Develop and scale managed inference services, including multi‑GPU, multi‑node parallelism and confidential‑compute capabilities.
  • Implement observability across control plane, GPU fleet and applications using metrics, logging and SLOs, and run an on‑call rotation.
  • Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests.
  • Collaborate with internal research, model‑training and product teams to align platform capabilities with workload needs.
  • Recruit and set technical standards for new engineers joining the platform team.

Required profile

  • 8+ years of hands‑on engineering experience, with at least 3 years leading teams that build and operate critical infrastructure platforms.
  • Bachelor’s or Master’s degree in Computer Science, Engineering or equivalent practical experience.
  • Proven experience running Slurm at scale for real users, including partitions, QoS, accounting and node‑health automation.
  • Deep knowledge of GPU fleet operation on bare metal: NVIDIA driver, CUDA lifecycle, MIG, DCGM health monitoring and node burn‑in.
  • Expertise in high‑performance interconnects such as InfiniBand, RDMA and SR‑IOV, and ability to troubleshoot multi‑node NCCL performance.
  • Strong Linux systems background, including kernel modules, PCIe passthrough, cgroups, namespaces and performance tuning.
  • Production‑grade Kubernetes operations experience, covering control‑plane upgrades, CNI/CSI, operators and multi‑tenancy design.
  • Hands‑on experience with HPC storage solutions (Lustre, NFS, VAST) and large‑scale data movement.
  • Fluency in JavaScript/Node.js sufficient to review and influence control‑plane services.

Required skills

  • Slurm
  • Kubernetes
  • Linux system administration
  • NVIDIA driver & CUDA
  • InfiniBand, RDMA, SR‑IOV
  • Prometheus, Grafana, Loki (observability)
  • JavaScript, Node.js
  • HPC storage (Lustre, NFS, VAST)
  • GPU fleet management (DCGM, MIG)

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Tether.io.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 8 hours ago

Expires 1 month from now

9 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Tether.io