We are looking for a highly motivated Machine Learning Evaluation Engineer to join our team and help define and drive the evaluation of advanced machine learning and computer vision technologies. You will work closely with algorithm, data, and engineering teams to build scalable evaluation methodologies, uncover model weaknesses, and turn complex data into actionable insights.
You will play a key role in ensuring that our ML systems deliver high-quality, robust experiences across diverse real-world scenarios. The ideal candidate combines strong fundamentals in classical machine learning and computer vision with hands-on experience in metrics design, data analysis, failure analysis, visualization, and evaluation infrastructure. You should also be comfortable leveraging modern AI-assisted tools and workflows to improve engineering efficiency and accelerate analysis.
In this role, you will define and drive the evaluation strategy for computer vision and machine learning algorithms used in complex product experiences. You will work closely with algorithm engineers to understand system behavior, identify the most meaningful quality signals, and develop evaluation frameworks that reflect real-world performance.
You will design metrics and evaluation methodologies that go beyond aggregate accuracy and help the team understand performance across important data slices and scenarios. You will analyze large-scale datasets to identify gaps in data quality and coverage and develop strategies to improve the representativeness of training and evaluation data.
A significant part of this role will involve failure analysis. You will investigate model failures, identify recurring patterns, develop failure taxonomies, and determine whether issues are driven by data, labeling, algorithm limitations, environmental conditions, or other system-level factors. You will build tools and visualizations that enable engineers to efficiently explore failures and understand the underlying root causes.
You will also develop scalable and automated workflows for evaluation, analysis, and reporting. You will be expected to leverage modern AI-assisted tools where appropriate to accelerate data analysis, visualization development, coding, and workflow automation while maintaining technical rigor and reproducibility.
As a member of the team, you will help shape evaluation best practices, influence algorithm and data decisions, and drive improvements across the ML development lifecycle. You should be comfortable navigating ambiguity, independently identifying opportunities for improvement, and partnering with cross-functional teams to deliver high-quality ML systems.