Research Evaluation Lead
Rhoda AI · San Francisco
OtherResearchFullTime
Rhoda AI is hiring a Research Evaluation Lead in San Francisco, California (full-time). Posted 11 September 2026. Apply directly on Rhoda AI's own careers site — no account needed.
At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.
Mission
Own end-to-end robot model evaluation for Research. Turn research questions into consistent, high-quality, repeatable evals that enable fast iteration and trusted results.
Responsibilities
Translate research intent into clear eval protocols, trial plans, and success criteria.
Own execution end-to-end: model handoff → station readiness → pilot execution → QA → results.
Train and manage eval pilots; ensure consistency across people, shifts, and stations.
Distinguish model failures from hardware, setup, operator, or data-quality issues.
Maintain eval setups, resets, randomization, metadata, and experiment traceability.
Track quality, throughput, and bottlenecks; continuously improve the eval process.
Partner closely with Research, Robot Data, and Eval Platform teams.
What we’re looking for
Some understanding of robotics / ML experimentation.
Computer science background or hands-on experience with coding
Rigorous, detail-oriented, and able to understand the intent behind an experiment, not just execute instructions.
Strong hands-on execution and ownership.
Experience in robotics testing, data collection, lab operations, or QA preferred.
Success looks like
A researcher can hand off a model and research question and receive a trusted, standardized eval result with sufficient trials and QA.
Listings are read directly from each employer's careers system (Greenhouse, Lever, Ashby, SmartRecruiters and Workday) and refreshed daily. hire.run never copies other job boards, and Apply always goes to the employer's own site.