Manager - Model Evaluation
TECHNICAL STACK · 2 TAGS
OVERVIEW
As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle (AV) prediction and planning stacks.
You will own the statistical frameworks, offline/online evaluation metrics, and validation pipelines that ensure our behavioral models operate safely, comfortably, and predictably. Operating at the intersection of Data Science, Machine Learning, and Safety Engineering, you will partner closely with Autonomy Software, Prediction, Planner, and ML Operations teams to establish data driven release criteria for our AV fleet.
IN THIS ROLE, YOU WILL
- Team Leadership & Execution: Lead, mentor, and scale a high-performing team of Data Scientists, ML Validation Engineers, and Software Engineers while driving roadmaps, sprint execution, resource allocation, and high-throughput model releases with rigorous safety guardrails. Culture of Rigor: Foster a culture of statistical excellence, healthy skepticism, proactive risk tracking, and data-driven decision-making.
- Validation Strategy & Methodologies: Define and execute end-to-end validation strategies across offline evaluation, open/closed-loop simulation, and shadow-mode fleet benchmarking to ensure robust behavioral model performance. Statistical uncertainties, and regressions into clear, data-driven recommendations for release gating and executive leadership.
- Metrics, Release Gating & Rigor: Oversee metric development and standardization with System Safety and Autonomy teams, establishing quantitative go/no-go release criteria for Behavioral Planner and Prediction ML models while fostering statistical rigor and proactive risk management.
- Cross-Functional & Infrastructure Partnership: Partner closely with Planner, Prediction, MLOps, and Developer Efficiency teams to translate behavioral requirements into measurable validation targets, streamline dataset and evaluation pipelines, and optimize runtime and compute costs.
- Executive Communication & Decision-Making: Translate complex model performance trade-offs, statistical uncertainty, regressions, and safety risks into clear, data-driven recommendations for release decisions and executive leadership.
QUALIFICATIONS
- Experience: Masters or PhD in CS, Robotics, Applied Statistics or a related field and 3+ years of direct engineering management experience leading Data Science, Machine Learning, or V&V engineering teams, alongside 7+ years of technical experience in robotics, autonomous systems, or AI/ML.
- Domain Knowledge: Strong background in ML model validation, behavioral evaluation frameworks, system-level performance benchmarking, and statistics.
- Software & Systems Literacy: Strong technical foundation in Python and modern data/ML platforms, with exposure to or conceptual literacy in large-scale production codebases (C++ or distributed systems). Proven ability to partner with systems software engineers, review technical architecture, and understand compute/performance trade-offs, Track record of leading teams evaluating complex robotic systems
- Technical Depth: Proven familiarity with modern C++/Python ML environments, simulation frameworks, high-throughput ML evaluation pipelines.
- Cross-Functional Leadership: Demonstrated ability to navigate complex organizational trade offs between release velocity, compute cost, and safety rigor.
BONUS QUALIFICATION
- Experience with autonomous vehicles, robotics, or other safety-critical systems.
- Experience building large-scale simulation, model evaluation, or validation infrastructure.
- Experience with reinforcement learning, generative AI, or distributed ML systems.
REQUIREMENTS
A Final Note
You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.
QUESTIONS AND ANSWERS
- How much does the Manager - Model Evaluation at Zoox pay?
- The posting lists a range of $237K–$338K per year. Ranges reflect what Zoox publicly declared on the source posting.
- Where is this Manager - Model Evaluation role based?
- The role is based in Foster City, CA.
- What experience does Zoox expect for this role?
- The posting is tagged as a executive-level role, typically 10+ years of experience. Check the requirements section for specifics.
- Where is Zoox headquartered?
- Zoox is headquartered in Foster City, USA.
- How was this posting sourced?
- This role was pulled directly from Zoox's Lever careers site. Apply links open in the employer's own ATS — no reposts or aggregator middleware.
Apply links open in the employer's official ATS. Always verify recruitment messages on the company's careers page before sharing personal information.