Product Pulse
Research Engineer, Training and Environment Infrastructure
Publicado 29 sept 2026
Referencia de mercado: mediana 108.750 €
basado en 24 ofertas en US
Este empleo está publicado en EN
About the company
We work on the data and infrastructure behind frontier model training and deployed AI agents. We partner
with AI labs and large organizations on post-training research, infrastructure, and deployments. We are an
early-stage, revenue-generating team hiring for in-person roles in San Francisco.
About the role
You sit at the center of the post-training loop. Before compute is spent on a dataset, you decide whether it
is worth it. After a model change ships, you determine whether it helped.
The work spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic
and task shapes evolve. This is not a fixed benchmark you maintain once and forget.
What you'll do
Rank training value before spending compute
Work out which tasks in a dataset are worth training on, and build a system that ranks every task by
expected training value.
Build scoring that runs cheaply at scale
Design automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good
outcome rather than generic correctness.
Research Engineer, RL Environments and Infrastructure
Catch drift before it becomes a problem
Notice when real customer requests have moved far enough from the existing test set that it no longer
describes the job, then rebuild the evaluations to match.
What we're looking for
- You have built environments, evaluations, or RL pipelines that ran in production.
- You have built RL data or done hands-on RL research, rather than working adjacent to it.
- You are a strong systems engineer who can pick up any stack.
- You have substantive opinions on reward design, including what makes a checker trustworthy and how
models learn to game them.
- You are comfortable owning ambiguous evaluation problems end to end, from building the first version
and inspecting the data to iterating until the system is useful.
- You can work six days a week, in person, in San Francisco.
Nice to have
- Experience with LLM-as-judge systems, automated graders, or rubric design for model outputs.
- A background in applied ML research, research engineering, or data science at a lab or fast-moving
startup.
- Experience building evaluation infrastructure tied directly to a training pipeline, not just offline
benchmarking.
Compensation and benefits
$150,000 to $350,000 USD base, depending on experience and seniority
- Meaningful equity grants for early team members
- Health, dental, and vision coverage
Fill in the form, we will contact you...
Resumen del puesto
Tipo de empleo
Tiempo completo
Correo electrónico
support@workfully.com
Habilidades requeridas
Empleos similares
Let's Get to Work
True North Recruiters
Remote Junior Sales Representative (Entry-Level)
Stratford Davis
Remote Customer Sales Representative
Stratford Davis
Virtual Sales Agent (Remote & Full Training Provided)
Stratford Davis
Evals Researcher, Rewards and Environments
Product Pulse
Full-Stack Product Engineer
Product Pulse