About this role
Role Overview
Review AI-assisted software development traces that support the training and evaluation of advanced AI models. Assess complete coding sessions created with AI-enabled developer tools, including code correctness, workflow quality, reasoning, and adherence to sound engineering practices. Deliver clear written feedback using established evaluation criteria.
Key Responsibilities- Evaluate the quality and correctness of end-to-end AI-assisted coding sessions.
- Judge coding workflows, technical reasoning, and implementation decisions for accuracy and best-practice alignment.
- Provide concise, rubric-based written feedback on reviewed traces.
- At least 3 years of professional software development experience.
- Hands-on experience with AI-assisted coding tools and agentic or specification-driven workflows, such as Cursor, GitHub Copilot, Claude Code, or similar tools.
- Strong code-reading and debugging abilities across full-stack or backend systems.
- Ability to assess multi-step coding trajectories for correctness and engineering best practices.
- Experience with Kiro or Amazon CodeCatalyst.
- Previous experience evaluating or grading AI-generated code.
- Contributions to developer tooling.
- Remote role available to candidates located in the United States.
- Hourly engagement.
$70 to $90 per hour.