Key details
- Role type
- Contract
- Compensation
- $29–$45/hr
- Work arrangement
- Remote
- Category
- technology
- Confirmed requirements
- 4
About this role
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and creating actionable red team data. This text-based remote role focuses on testing AI responses for issues including bias, misinformation, harmful behavior, and misuse risks.
Key Responsibilities- Red team conversational AI models and agents through jailbreak attempts, prompt injection, misuse scenarios, bias exploitation, and multi-turn manipulation.
- Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors. Higher-sensitivity assignments are optional, supported by clear guidelines and wellness resources, and topics are communicated before exposure.
- Annotate failures, classify vulnerabilities, and flag systemic risks.
- Use testing taxonomies, benchmarks, and playbooks to apply consistent evaluation methods.
- Create reproducible reports, datasets, and attack cases that support improvements to AI systems.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- A curious, adversarial approach to identifying system breaking points.
- Ability to use structured frameworks or benchmarks for testing.
- Ability to explain risks clearly to both technical and non-technical stakeholders.
- Adaptability across projects and customer environments.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk testing, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Psychology, acting, or writing used for unconventional adversarial thinking.
- Remote, hourly engagement.
- Compensation: $29 to $45 per hour.
- Native fluency in English and Portuguese is required.
- Portuguese fluency must be in a global Portuguese variant other than Brazilian Portuguese.
What to prepare before applying
- Remote work location.
- Native fluency in English and Portuguese (global, excluding Brazilian Portuguese).
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to use frameworks or benchmarks for structured testing.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.