All roles
GPU Infrastructure SRE
Own the reliability and operations of large-scale GPU infrastructure for model training and production inference.
About the Role
- Build and operate the infrastructure that keeps large-model training and online inference running reliably.
- You will support GPU clusters at 100+ card scale, respond to production incidents, and work closely with training and inference teams on stability, performance, and resource efficiency.
- This is a hands-on infrastructure role across Linux, Slurm, Kubernetes, Ceph, RDMA networking, GPU drivers, observability, and automation.
Responsibilities
- Provide infrastructure support for large-model training and online inference workloads, responding quickly to production incidents and operational issues.
- Deploy and operate 100+ GPU clusters, ensuring training and inference jobs run stably at scale.
- Maintain Slurm scheduling and Kubernetes platforms, optimizing resource allocation and multi-tenant isolation.
- Deploy, expand, and tune Ceph distributed storage for large-scale AI workloads.
- Operate RDMA networks such as InfiniBand and RoCE, plus GPU fleet components including DCGM, drivers, CUDA, and NCCL version management.
- Build automation tools, monitoring, and alerting systems that improve cluster stability and reduce debugging time.
Requirements
- Senior Linux operations background, with hands-on experience in GPU clusters or HPC environments.
- Experience operating 100+ GPU clusters that support large-model training or online inference workloads.
- Deep understanding of the Linux kernel, networking stack, and storage stack, with the ability to diagnose low-level performance bottlenecks.
- Expertise with Slurm for training workloads and production Kubernetes clusters for inference, including architecture design and performance tuning.
- Strong experience with Ceph distributed storage, including large-scale deployment, expansion, and performance tuning.
- Familiarity with RDMA networking such as InfiniBand or RoCE, and the ability to help debug NCCL collective communication issues.
- Strong Python and Shell scripting skills, with experience building automation platforms or operations tooling.
- Experience with Prometheus and Grafana monitoring systems, including large-scale cluster monitoring and alerting.
- Strong ownership for production reliability, including leading incident response, postmortems, and on-call responsibilities.
Nice to Have
- Ability to partner with training teams to diagnose performance bottlenecks across NCCL communication and storage I/O.
- Experience building GPU clusters from zero to one at 100+ to 1,000+ GPU scale.
- Familiarity with additional distributed storage systems such as MinIO, Weka, or Lustre.
- HPC operations experience and familiarity with parallel file systems.
- Open-source contributions or technical writing related to infrastructure, HPC, GPU operations, or reliability.
How to Apply
- Send your resume along with GitHub, personal project, or technical writing links to the contact below.
- For open-source contributions or past projects, direct links or short write-ups are welcome.
- Take-home tasks, if any, will be paid at a reasonable market rate.
- No requirements around years of experience or degree - we evaluate on technical depth and past work.
Explore More
Other Open Roles
Brand and Creative Designer
Own visual direction and creative execution across campaigns, launches, social media, product visuals, and the long-term brand system.
View roleVoice Acquisition & Quality Lead
Own the quality bar, acquisition pipeline, and catalog curation for natural-sounding voices on Fish Audio.
View roleAI Systems Engineer
Design and optimize high-performance algorithm engines for in-house large model development and real-world AI applications.
View roleVoice Agent Platform Engineer
Build the architecture, engine, SDKs, and product surfaces for Fish Audio's open realtime voice Agent platform.
View role