Applied AI Research Engineer

Beam · New York, NY, US | San Francisco, CA, US · Posted 2mo ago

$140k - $200k

AI-Native Cloud Platform

**Beam** is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

# **About the Role**

We’re looking to hire someone to own inference research hands-on and find ways to lower cost per token and latency on our customer workloads.

* Low-level inference optimization, from speculative decoding, quantization, KV-cache and memory management * Work directly with customers to optimize their production workloads, and apply your learnings to our platform as product improvements * High-level of autonomy to find the highest upside bets and guide the future of our inference platform based on your work

# Skills & Experience

* Systems or research background in LLM inference * Deep understanding of LLM serving, from the kernel to the scheduler * History of shipping products or research that people use in production-like scenarios, whether academic or industry * Excited to collaborate closely with customers * Enthusiasm for developer tools, cloud native technologies, and open source software

# Benefits

* Competitive salary and meaningful equity * Join a fast-growing pre-series A company at the ground floor * Health, dental, and vision benefits with 90% coverage for you and 50% for dependents * Opportunities to participate in events across the cloud native community * Fitness stipend, learning budget, and much, much more

Apply on Beam's site