Research Engineer (General)

HUD · San Francisco | Singapore · In-office · Posted 2d ago

$140k - $250k

Platform for building RL environments and evals

## **About HUD**

[HUD](https://www.hud.ai/) is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.

## **About the role**

_This is a general application for candidates who are unsure which research focus -_ [_QC Automation_](https://jobs.ashbyhq.com/hud/e6f9812e-dcfd-422b-b614-1d1273c16003)_,_ [_Benchmarks_](https://jobs.ashbyhq.com/hud/c5252f0e-fd1d-41cf-b803-b0ce1fe4cea1)_, or_ [_Synthetic Data_](https://jobs.ashbyhq.com/hud/44e356fa-801c-4dac-99f5-848d80a68500) _- they would be a fit for. We would love to meet you and figure it out together. However, if you already have a focus in mind, **please apply to only that application**._

We're looking for Research Engineers to build the technical foundation for training and evaluating frontier AI agents. You’ll build the systems for creating new environments, improve data quality, and translate real-world workflows into tasks and benchmarks.

## **Responsibilities**

- Build systems for creating, running, evaluating, and improving agent training environments

- Design experiments to understand model behavior, agent failure modes, and data quality issues

- Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops

- Work across the full lifecycle of agent training data - task design, environment setup, trajectory collection, evaluation, and validation

- Partner with external vendors to identify bottlenecks and improve the quality and throughput of HUD’s data engine

- Build metrics and analyses that help us understand whether our tasks, environments, and evals are actually useful for training frontier agents

## **Experience**

**You may be a good fit if you have:**

- Proficiency in Python, Docker, and Linux environments

- Experience working on benchmarks and evals - you can reason about what makes a task realistic, a rubric reliable, an environment usable, and a trajectory useful for RL training

- Strong attention to detail and the ability to spot subtle inconsistencies in data, model behavior, or task design

- Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap

- Early-stage startup experience with ability to work independently in fast-paced environments

**Strong candidates may also have:**

- Experience building internal tools, research infrastructure, or data pipelines

- Experience designing metrics and validation workflows

- A background in competitive programming, Olympiad medaling, research, or unusually strong independent project experience

- Thrive in unstructured problem spaces

- Strong communication skills for remote collaboration across time zones

_We prioritize technical aptitude and learning potential over years of experience. Motivated candidates are encouraged to apply even if they don't meet all criteria._

## **Team & company details**

- **Team Size** : ~25 people currently, mostly full-time in-person, but some remote.

- **Our team:** Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.

- **Company stage:** We have 8 figures in funding and are scaling profitably and quickly to meet very strong demand.

## **Logistics**

- **Employment** : Full-time.

- **Location** : We have offices in San Francisco or Singapore but are open to remote candidates who **can work hours that 70-80% overlap with either San Francisco or Singapore time zones.**

- **Visa Sponsorship** : We provide support for relocation and visas for strong full-time candidates to the US or Singapore.

- **Timeline** : Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.

## What we offer

- Competitive compensation

- 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)

- Lunch and dinner when you’re in the office (in-office employees)

- Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays

- Other perks including an Equinox membership, 401k, and commuter benefits (US employees)

- Unlimited\* access to tokens for ChatGPT, Claude Code, Cursor, etc. \*_By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is._

Due to high volume, we may not actively respond to every application, but feel free to contact us at [[email removed]](mailto:[email removed]) or elsewhere if we missed your application!

Apply on HUD's site