Product Pulse
Machine Learning Engineer – Evals
Publicado 6 oct 2026
Este empleo está publicado en EN
Machine Learning Engineer – Evals
📍 New York, NY (on-site, 5 days a week) | 💰 $220K–$300K + competitive equity | Full-time
Build the measuring stick for AI memory and identity.
Our client, a research-driven AI startup, builds the memory and identity layer for the agentic world. Their product works, and it's growing more sophisticated by the week. It's a multi-agent system with real black-box behaviour, and the team needs one person to answer a hard question: are our identity representations actually getting better?
There's no ground truth to lean on. You'll define what "better" means for representations that change over time, then build the machinery that measures it. If you enjoy constructing measurements from nothing, digging through traces to find what's really broken, and shipping the fix yourself, this role is built for you.
What you'll do
- Design the evals. Define the scores and the meaning of "better" for an evolving entity representation, and keep that definition current as the product and methods move.
- Build the pipelines and harnesses. Data ingestion, labelling, versioning, reruns and LLM judges: the infrastructure that lets the team ask a new question this week and get an answer this week.
- Run them and find the insight. Read traces and results, surface what's actually broken (not just what's easy to measure) and propose fixes that lead to higher-fidelity representations.
- Build simulation agents at scale. Iteration speed is capped by how many entities can be modelled and measured at once. You'll raise that ceiling.
- Own the loop end to end. The question, the pipeline, the rerun, the write-up. No handoffs, no waiting on someone else's sprint.
What you'll bring
- A track record of designing, building and running LLM evals end to end
- 3+ years designing and owning LLM evaluation
- Production-quality Python: tests, maintenance and a steady release cadence
- Hands-on experience with PyTorch, Hugging Face and Weights & Biases
- Data pipeline tooling (Hydra, Kafka, SQL) is a plus
Bonus points for
- Work on memory, identity, personalization, agent evaluation or agentic optimization
- Shipped research: first or co-author at NeurIPS, ICML, ICLR or equivalent, or open-source work of comparable quality
- A degree in computer science, mathematics, statistics or a related quantitative field
Good to know
- This role is fully on-site in New York, five days a week.
- Candidates must be authorized to work in the US without visa sponsorship (US citizens or green card holders).
- One seat available.
Interested? Apply here.
Talent Finest connects exceptional talent with high-growth startups and leading organizations across the US, Europe and Africa.
Fill in the form, we will contact you...
Resumen del puesto
Tipo de empleo
Tiempo completo
Correo electrónico
support@workfully.com
Habilidades requeridas
Empleos similares
Software Engineer
Product Pulse
Senior Software Engineer (Full Stack)
Product Pulse
Part-Time PM
Shyft6
Full-Stack Software Engineer
Qvid Talent Solutions
Ausbildung Bankkauffrau (m/w/d)
LBS Bezirksdirektor Michael Scheffner - LBS Südwest Beratungsstelle Montabaur
Ausbildung Bankkauffrau (m/w/d)
LBS Bezirksdirektor Michael Scheffner - LBS Südwest Beratungsstelle Montabaur