Site Reliability Engineer

Beam · New York, NY, US | San Francisco, CA, US · Posted 2mo ago

$140k - $200k

AI-Native Cloud Platform

**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

# **About the Role**

* Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production. * Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage. * Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production. * Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.

# Skills & Experience

* You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it * You're fluent with AI tooling. You aren’t afraid to max-out your token usage for the right spec. * You’re comfortable debugging production issues, from triage to post-mortem. * Enthusiasm for developer tools, cloud native technologies, and open source software

# Benefits

* Competitive salary and meaningful equity * Join a fast-growing pre-series A company at the ground floor * Health, dental, and vision benefits with 90% coverage for you and 50% for dependents * Opportunities to participate in events across the cloud native community * Fitness stipend, learning budget, and much, much more

Apply on Beam's site