← Back to your search

STN Inc

Site Reliability Engineer

DevOps / SRE

Employment

Full-time
From the employer

FullTime

Level

Not stated

Team

AI / AI
Not doable from Austria

Country assessment

Not doable from Austria because it is remote but scoped to United States.

Our assessment is guidance. Confirm arrangements with the employer.

Skills mentioned in this posting

JavaScriptPythonGoKubernetesTerraformLinuxObservability

Job description

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer keeps GPU One (GPUaaS) reliable, observable, and automated at scale. This role blends deep systems troubleshooting with strong software automation to eliminate toil and ensure the GPU compute platforms that run customer AI workloads stay fast and available. It is a hands-on position: you will work with root access on production head and compute nodes, and you should be as comfortable at an Ubuntu command line as you are in an editor.

Key responsibilities

  • Own the reliability, availability, and performance of production Linux GPU clusters, from the operating system up through drivers, GPUs, high-speed networking, and storage

  • Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues

  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths

  • Diagnose network faults end to end — configuration, routing, and the physical layer, including cabling, transceivers, and link errors

  • Configure and maintain the workload managers that schedule customer jobs — Kubernetes, Slurm, or both — along with the identity, storage, and networking services they depend on

  • Build automation and tooling that eliminates operational toil, with enough fluency in a modern language such as Python or Go to design the solution, review what is written, and redirect an approach that is heading the wrong way

  • Use AI-assisted engineering tools such as Claude to accelerate automation, runbooks, and incident analysis

  • Automate provisioning, image deployment, configuration, and remediation with Ansible and infrastructure-as-code

  • Design and operate observability with Grafana, Prometheus, and Loki — tuning alerts toward signal and building self-healing that takes a human out of the loop

  • Lead incident response, on-call, blameless postmortems, and the reliability improvements that keep us inside customer SLAs

  • Partner with Platform and Systems Engineering on capacity, rollout, and continuous improvementRequired qualifications

  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications

  • 5+ years in systems, infrastructure, or SRE engineering operating production systems at scale

  • Deep Linux troubleshooting skills across OS, networking, storage, and performance, with real comfort working as root on production systems (we run Ubuntu almost exclusively)

  • Hands-on experience operating GPU servers in production — driver, device, and hardware-level problems, not only the workloads running on top of them

  • Practical network troubleshooting, including physical-layer faultsStrong automation instincts and the programming ability behind them, in Python or a comparable language

  • Experience with configuration management, node provisioning, and IaC using Ansible, Terraform, or similar

  • Experience building observability and alerting with Grafana and Prometheus

  • Bachelor’s degree in computer science or equivalent experience