Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Red Team Expert, English and Portuguese

$29–$45/hr

RemoteContracttechnology
Apply Now

Key details

Role type
Contract
Compensation
$29–$45/hr
Work arrangement
Remote
Category
technology
Confirmed requirements
4

About this role

Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and creating actionable red team data. This text-based remote role focuses on testing AI responses for issues including bias, misinformation, harmful behavior, and misuse risks.

Key Responsibilities
  • Red team conversational AI models and agents through jailbreak attempts, prompt injection, misuse scenarios, bias exploitation, and multi-turn manipulation.
  • Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors. Higher-sensitivity assignments are optional, supported by clear guidelines and wellness resources, and topics are communicated before exposure.
  • Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Use testing taxonomies, benchmarks, and playbooks to apply consistent evaluation methods.
  • Create reproducible reports, datasets, and attack cases that support improvements to AI systems.
Qualifications
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • A curious, adversarial approach to identifying system breaking points.
  • Ability to use structured frameworks or benchmarks for testing.
  • Ability to explain risks clearly to both technical and non-technical stakeholders.
  • Adaptability across projects and customer environments.
Preferred Specialties
  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
  • Cybersecurity, including penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk testing, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
  • Psychology, acting, or writing used for unconventional adversarial thinking.
Work Terms
  • Remote, hourly engagement.
  • Compensation: $29 to $45 per hour.
Eligibility
  • Native fluency in English and Portuguese is required.
  • Portuguese fluency must be in a global Portuguese variant other than Brazilian Portuguese.

What to prepare before applying

  1. Remote work location.
  2. Native fluency in English and Portuguese (global, excluding Brazilian Portuguese).
  3. Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  4. Ability to use frameworks or benchmarks for structured testing.

These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.

Related Jobs

More like this