Technical Lead - GPU Infrastructure
Tether.io
Job description
About the role
Join Tether Data's Cosmic AC GPU compute and managed inference platform as Technical Lead. You will own the architecture and delivery of a full‑stack bare‑metal GPU infrastructure, leading a distributed team of about twelve engineers across backend, frontend, DevOps, QA and documentation.
Key responsibilities
- Own the end‑to‑end platform architecture, producing high‑level and low‑level designs and keeping them up to date.
- Lead and line‑manage a distributed team, establishing engineering standards, conducting code reviews, and driving growth.
- Design, build and operate a managed Slurm scheduling layer for research and model‑training workloads.
- Bootstrap and maintain the Kubernetes control plane on partner‑provided bare‑metal, handling GPU operator, network operator, upgrades and day‑2 operations.
- Develop and scale managed inference services, including multi‑GPU, multi‑node parallelism and confidential‑compute capabilities.
- Implement observability across control plane, GPU fleet and applications using metrics, logging and SLOs, and run an on‑call rotation.
- Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests.
- Collaborate with internal research, model‑training and product teams to align platform capabilities with workload needs.
- Recruit and set technical standards for new engineers joining the platform team.
Required profile
- 8+ years of hands‑on engineering experience, with at least 3 years leading teams that build and operate critical infrastructure platforms.
- Bachelor’s or Master’s degree in Computer Science, Engineering or equivalent practical experience.
- Proven experience running Slurm at scale for real users, including partitions, QoS, accounting and node‑health automation.
- Deep knowledge of GPU fleet operation on bare metal: NVIDIA driver, CUDA lifecycle, MIG, DCGM health monitoring and node burn‑in.
- Expertise in high‑performance interconnects such as InfiniBand, RDMA and SR‑IOV, and ability to troubleshoot multi‑node NCCL performance.
- Strong Linux systems background, including kernel modules, PCIe passthrough, cgroups, namespaces and performance tuning.
- Production‑grade Kubernetes operations experience, covering control‑plane upgrades, CNI/CSI, operators and multi‑tenancy design.
- Hands‑on experience with HPC storage solutions (Lustre, NFS, VAST) and large‑scale data movement.
- Fluency in JavaScript/Node.js sufficient to review and influence control‑plane services.
Required skills
- Slurm
- Kubernetes
- Linux system administration
- NVIDIA driver & CUDA
- InfiniBand, RDMA, SR‑IOV
- Prometheus, Grafana, Loki (observability)
- JavaScript, Node.js
- HPC storage (Lustre, NFS, VAST)
- GPU fleet management (DCGM, MIG)
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Egypt.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 8 hours ago
Expires 1 month from now
9 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Tether.io