All roles

Platform

SRE / Platform Engineer, GPU Infrastructure

About the role

Based in San Mateo, California, you will build and operate a large-scale GPU compute platform and remotely maintain the GPU cluster at our Santa Clara data center.

This is a highly hands-on role spanning Linux, Kubernetes, networking, automation, observability, and GPU hardware. You will work closely with research and engineering teams to make compute reliable, easy to use, and ready for rapid iteration. The work ranges from troubleshooting Kubernetes workloads and Linux networking to deploying bare-metal servers and diagnosing performance bottlenecks in multi-node GPU systems.

Your regular workplace will be our San Mateo office, with visits to the data center for hands-on equipment work as needed.

What you’ll do

  1. Deploy, operate, expand, and troubleshoot bare-metal NVIDIA and AMD GPU clusters.
  2. Build and maintain Kubernetes platforms for training, inference, and research computing.
  3. Manage Linux, container runtimes, NVIDIA drivers, CUDA, and the GPU Operator.
  4. Configure and troubleshoot VLANs, routing, DNS, firewalls, and high-speed networks.
  5. Automate server provisioning, configuration, upgrades, monitoring, and recovery.
  6. Build monitoring and alerting for GPUs, servers, networks, storage, and Kubernetes.
  7. Improve resource scheduling, GPU utilization, platform reliability, and isolation between users.
  8. Build internal tools that help research teams access compute and troubleshoot their workloads.
  9. Perform hands-on rack installation, cabling, BMC/IPMI management, and hardware troubleshooting.
  10. Participate in on-call support for critical infrastructure and contribute to incident reviews.

What you’ll bring

  • Experience with Linux administration, performance analysis, and troubleshooting.
  • Practical experience operating Kubernetes and container infrastructure.
  • A working understanding of TCP/IP, VLANs, routing, DNS, NAT, and firewalls.
  • Experience automating infrastructure with tools such as Go, Bash, Ansible, or Terraform.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, or Loki.
  • The ability to troubleshoot across hardware, operating systems, networks, containers, and applications.
  • Willingness to work on-site in San Mateo and visit the data center for hands-on equipment work when needed.

Nice to have

  • Experience with NVIDIA or AMD GPU clusters.
  • Familiarity with NCCL, NVLink, RDMA, InfiniBand, RoCE, or GPUDirect.
  • Experience deploying or troubleshooting 100G–800G networks.
  • Experience with schedulers and distributed compute tools such as Slurm, Ray, KubeRay, or Kueue.
  • Experience with bare-metal lifecycle management using PXE, Redfish, or IPMI.
  • Familiarity with inference systems such as vLLM, SGLang, or Triton.
  • Experience with virtualization technologies such as KVM, Proxmox, KubeVirt, Incus, or Firecracker.
  • Experience with Ceph, ZFS, NVMe, or distributed storage.
  • Experience with eBPF, the Linux kernel, or performance optimization.
  • Familiarity with data center power, cooling, rack planning, and structured cabling.

How to apply

Email careers@bakelab.ai to introduce yourself and share your relevant experience.