All roles

GPU Infrastructure SRE

Own the reliability and operations of large-scale GPU infrastructure for model training and production inference.

About the Role

  • Build and operate the infrastructure that keeps large-model training and online inference running reliably.
  • You will support GPU clusters at 100+ card scale, respond to production incidents, and work closely with training and inference teams on stability, performance, and resource efficiency.
  • This is a hands-on infrastructure role across Linux, Slurm, Kubernetes, Ceph, RDMA networking, GPU drivers, observability, and automation.

Responsibilities

  • Provide infrastructure support for large-model training and online inference workloads, responding quickly to production incidents and operational issues.
  • Deploy and operate 100+ GPU clusters, ensuring training and inference jobs run stably at scale.
  • Maintain Slurm scheduling and Kubernetes platforms, optimizing resource allocation and multi-tenant isolation.
  • Deploy, expand, and tune Ceph distributed storage for large-scale AI workloads.
  • Operate RDMA networks such as InfiniBand and RoCE, plus GPU fleet components including DCGM, drivers, CUDA, and NCCL version management.
  • Build automation tools, monitoring, and alerting systems that improve cluster stability and reduce debugging time.

Requirements

  • Senior Linux operations background, with hands-on experience in GPU clusters or HPC environments.
  • Experience operating 100+ GPU clusters that support large-model training or online inference workloads.
  • Deep understanding of the Linux kernel, networking stack, and storage stack, with the ability to diagnose low-level performance bottlenecks.
  • Expertise with Slurm for training workloads and production Kubernetes clusters for inference, including architecture design and performance tuning.
  • Strong experience with Ceph distributed storage, including large-scale deployment, expansion, and performance tuning.
  • Familiarity with RDMA networking such as InfiniBand or RoCE, and the ability to help debug NCCL collective communication issues.
  • Strong Python and Shell scripting skills, with experience building automation platforms or operations tooling.
  • Experience with Prometheus and Grafana monitoring systems, including large-scale cluster monitoring and alerting.
  • Strong ownership for production reliability, including leading incident response, postmortems, and on-call responsibilities.

Nice to Have

  • Ability to partner with training teams to diagnose performance bottlenecks across NCCL communication and storage I/O.
  • Experience building GPU clusters from zero to one at 100+ to 1,000+ GPU scale.
  • Familiarity with additional distributed storage systems such as MinIO, Weka, or Lustre.
  • HPC operations experience and familiarity with parallel file systems.
  • Open-source contributions or technical writing related to infrastructure, HPC, GPU operations, or reliability.

How to Apply

  • Send your resume along with GitHub, personal project, or technical writing links to the contact below.
  • For open-source contributions or past projects, direct links or short write-ups are welcome.
  • Take-home tasks, if any, will be paid at a reasonable market rate.
  • No requirements around years of experience or degree - we evaluate on technical depth and past work.