About this role
Role Overview
Join a talent network of experienced data scientists who may contribute to future projects evaluating how effectively AI systems perform real-world data science work for leading AI research organizations.
Key Responsibilities
- Create precise, task-specific grading criteria for data science deliverables, including exploratory analyses, statistical modeling, machine learning pipelines, experimentation and A/B test write-ups, feature engineering, and technical reports or notebooks.
- Evaluate AI-generated or human-created work using established criteria.
- Provide detailed written explanations supporting evaluations and scores.
- Apply consistent, evidence-based judgment to produce reproducible, defensible assessments.
- Incorporate structured feedback from senior reviewers and revise submitted work accordingly.
Responsibilities will vary by project.
Qualifications
- At least 1 year of professional data science experience.
- Experience at a leading technology, research, or quantitative firm, such as a top FAANG company, AI lab, top-tier quantitative fund, or equivalent organization.
- Strong Python and SQL skills, plus expertise in statistical modeling, machine learning, experimentation, causal inference, and turning messy real-world data into rigorous analyses.
- Exceptional written communication skills and the ability to explain technical findings clearly.
- A detail-oriented, consistent approach to evaluating complex work.
- Comfort receiving feedback and calibrating judgment to established standards.
Work Terms
- Remote, hourly engagement.
- There is no immediate project opening. Qualified applicants may be contacted as relevant opportunities become available.
Compensation
$100 to $150 per hour.