Product Pulse
Research Engineer, Training and Environment Infrastructure
2026년 9월 29일에 게시됨
이 채용공고는 EN로 게시되어 있습니다
About the company
We work on the data and infrastructure behind frontier model training and deployed AI agents. We partner
with AI labs and large organizations on post-training research, infrastructure, and deployments. We are an
early-stage, revenue-generating team hiring for in-person roles in San Francisco.
About the role
You sit at the center of the post-training loop. Before compute is spent on a dataset, you decide whether it
is worth it. After a model change ships, you determine whether it helped.
The work spans reward design, automatic scoring at scale, and staying ahead of drift as customer traffic
and task shapes evolve. This is not a fixed benchmark you maintain once and forget.
What you'll do
Rank training value before spending compute
Work out which tasks in a dataset are worth training on, and build a system that ranks every task by
expected training value.
Build scoring that runs cheaply at scale
Design automatic scoring cheap enough to run constantly, tuned to each customer's definition of a good
outcome rather than generic correctness.
Research Engineer, RL Environments and Infrastructure
Catch drift before it becomes a problem
Notice when real customer requests have moved far enough from the existing test set that it no longer
describes the job, then rebuild the evaluations to match.
What we're looking for
- You have built environments, evaluations, or RL pipelines that ran in production.
- You have built RL data or done hands-on RL research, rather than working adjacent to it.
- You are a strong systems engineer who can pick up any stack.
- You have substantive opinions on reward design, including what makes a checker trustworthy and how
models learn to game them.
- You are comfortable owning ambiguous evaluation problems end to end, from building the first version
and inspecting the data to iterating until the system is useful.
- You can work six days a week, in person, in San Francisco.
Nice to have
- Experience with LLM-as-judge systems, automated graders, or rubric design for model outputs.
- A background in applied ML research, research engineering, or data science at a lab or fast-moving
startup.
- Experience building evaluation infrastructure tied directly to a training pipeline, not just offline
benchmarking.
Compensation and benefits
$150,000 to $350,000 USD base, depending on experience and seniority
- Meaningful equity grants for early team members
- Health, dental, and vision coverage
Fill in the form, we will contact you...
직무 요약
직무 유형
정규직
support@workfully.com
필수 기술
유사 채용공고
Let's Get to Work
True North Recruiters
Evals Researcher, Rewards and Environments
Product Pulse
Full-Stack Product Engineer
Product Pulse
Senior Growth Engineer
Product Pulse
Licensed Insurance Agent Benefits Enrollment
Employee Retention Benefits
Founding AI Engineer, Computer Vision
Product Pulse