STN Inc
Site Reliability Engineer
Employment
Full-timeFrom the employer
FullTime
Level
Not statedTeam
AI / AICountry assessment
Not doable from Austria because it is remote but scoped to United States.
Our assessment is guidance. Confirm arrangements with the employer.
Skills mentioned in this posting
Job description
Site Reliability Engineer
Platform and software · shared across customers
Reports to: Director, Site Reliability
Location: Remote (US)
Department: Cloud Platform Engineering / SRE/Reliability
Position summary
The Site Reliability Engineer keeps GPU One (GPUaaS) reliable, observable, and automated at scale. This role blends deep systems troubleshooting with strong software automation to eliminate toil and ensure the GPU compute platforms that run customer AI workloads stay fast and available. It is a hands-on position: you will work with root access on production head and compute nodes, and you should be as comfortable at an Ubuntu command line as you are in an editor.
Key responsibilities
Own the reliability, availability, and performance of production Linux GPU clusters, from the operating system up through drivers, GPUs, high-speed networking, and storage
Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues
Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths
Diagnose network faults end to end — configuration, routing, and the physical layer, including cabling, transceivers, and link errors
Configure and maintain the workload managers that schedule customer jobs — Kubernetes, Slurm, or both — along with the identity, storage, and networking services they depend on
Build automation and tooling that eliminates operational toil, with enough fluency in a modern language such as Python or Go to design the solution, review what is written, and redirect an approach that is heading the wrong way
Use AI-assisted engineering tools such as Claude to accelerate automation, runbooks, and incident analysis
Automate provisioning, image deployment, configuration, and remediation with Ansible and infrastructure-as-code
Design and operate observability with Grafana, Prometheus, and Loki — tuning alerts toward signal and building self-healing that takes a human out of the loop
Lead incident response, on-call, blameless postmortems, and the reliability improvements that keep us inside customer SLAs
Partner with Platform and Systems Engineering on capacity, rollout, and continuous improvementRequired qualifications
5+ years in SRE, DevOps, or production engineering roles
Strong programming skills in Go, Python, or both
Hands-on experience operating Kubernetes-based platforms at scale
Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
Strong incident management experience including major-incident command
Preferred qualifications
5+ years in systems, infrastructure, or SRE engineering operating production systems at scale
Deep Linux troubleshooting skills across OS, networking, storage, and performance, with real comfort working as root on production systems (we run Ubuntu almost exclusively)
Hands-on experience operating GPU servers in production — driver, device, and hardware-level problems, not only the workloads running on top of them
Practical network troubleshooting, including physical-layer faultsStrong automation instincts and the programming ability behind them, in Python or a comparable language
Experience with configuration management, node provisioning, and IaC using Ansible, Terraform, or similar
Experience building observability and alerting with Grafana and Prometheus
Bachelor’s degree in computer science or equivalent experience