Home Job Detail
About the Role You will own the design and implementation of rigorous, domain-specific benchmarks used to evaluate frontier AI agents on realistic workflows. Sitting within a small, highly technical team of researchers and engineers, this role is central to delivering evaluations that AI labs and enterprise customers genuinely trust and rely on. What You'll Do Design, implement, and maintain high-quality internal benchmarks for evaluating frontier agents on domain-specific tasks. Partner with s…