Home Job Detail
About the Role This is a hands-on research engineering role focused on designing and owning high-quality benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You will sit within a small, highly technical team and play a critical part in ensuring evaluations are rigorous, credible, and trusted by leading AI labs and customers. What You'll Do Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks. Partn…