Voltar às vagas
ScovaiScovaiJobs
Product Pulse

Product Pulse

Evals Researcher, Rewards and Environments

San Francisco, USPresencialContratoTempo integral

Publicado 6 de out. de 2026

Esta vaga foi publicada em EN

About the Role: Evals Researcher, Rewards and Environments, San Francisco (hybrid)

Post-training and evaluation lab, profitable at four months, San Francisco

You would build the evaluation systems that decide which data is worth training on and whether the models and agents are actually getting better.


About the Company

Our company builds the evaluation environments and post-training infrastructure that make long-horizon enterprise agents trainable, then uses that stack to train and deploy custom models for very large technology companies and government customers. Under four months in, ten people, already profitable, on angel and institutional money and no priced round.


What you would do:

- Rank training value before compute is spent: work out which tasks in a dataset are worth training on and build a system that ranks every task by expected training value.

- Build scoring that runs constantly and cheaply, tuned to each customer's definition of a good outcome rather than generic correctness.

- Catch drift before it is a problem: notice when real customer requests have moved far enough from the test set that it no longer describes the job, and rebuild the evals to match.

- Design rewards: checkers a model cannot game, model-as-judge and rubric graders, and evaluation of long-horizon agent trajectories.

- Sit at the centre of the post-training loop: decide before compute is spent, and say afterwards whether a change actually helped.




Hard requirements · from the JD

  • Built RL data or done RL research hands-on, not adjacent

  • Built an evaluation system themselves (graders, model-as-judge, environments or RL pipelines) that ran on production traffic or behind a published result

  • Real opinions on reward design: what makes a checker trustworthy and how models game it

  • Strong systems engineer who can pick up any stack; an engineering title on the current role

  • At least one year full-time, one to five typical

  • Six days of work a week, hybrid in San Francisco (three or four in the office)


Fill in the form, we will contact you...

Resumo da função

Tipo de vaga

Tempo integral

Email

support@workfully.com

Competências necessárias

Reinforcement Learning research (hands-on)RL data engineering (building and curating RL training datasets)Evaluation system implementation (graders, model-as-judge, environments, RL pipelines)Reward design and robustness against gaming (designing checkers and rubrics)Data valuation / training-value ranking (estimating which tasks are worth training on)Scoring and metric design aligned to customer-defined outcomesDrift detection and dataset shift monitoring (noticing when production requests diverge from test sets)Experimentation and causal evaluation (determining whether a change actually helped)Systems engineering and full-stack adaptability (able to pick up any stack)Production deployment and reliability for evaluation pipelines (running on production traffic)Long-horizon agent trajectory evaluationObservability and continuous scoring (cheap, constant evaluation runs and monitoring)

Vagas similares

True North Recruiters

Let's Get to Work

True North Recruiters

San Francisco, USPresencialContratoTempo integral
anteontem
Stratford Davis

Remote Junior Sales Representative (Entry-Level)

Stratford Davis

San Francisco, USPresencialContratoTempo integral
anteontem
Stratford Davis

Remote Customer Sales Representative

Stratford Davis

San Francisco, USPresencialContratoTempo integral
anteontem
Stratford Davis

Virtual Sales Agent (Remote & Full Training Provided)

Stratford Davis

San Francisco, USPresencialContratoTempo integral
anteontem
Product Pulse

Full-Stack Product Engineer

Product Pulse

San Francisco, USPresencialContratoTempo integral
há 4 dias
Product Pulse

Senior Growth Engineer

Product Pulse

San Francisco, USPresencialContratoTempo integral
há 4 dias