Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Red Team Expert, English & Danish

$48–$62/hr

RemoteContracttechnology
Apply Now

About this role

Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and producing practical safety data. This remote role focuses on text-based AI red teaming in both English and Danish.

Key Responsibilities
  • Test conversational AI models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
  • Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors.
  • Annotate failures, classify vulnerabilities, flag systemic risks, and create high-quality human data.
  • Use established taxonomies, benchmarks, and playbooks to apply testing consistently.
  • Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
Qualifications
  • Native fluency in both English and Danish is required.
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Ability to use frameworks and benchmarks to test systems methodically.
  • Clear communication skills for explaining risks to technical and non-technical audiences.
  • Comfort adapting across projects and customer contexts.
Preferred Experience
  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
  • Cybersecurity, including penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk testing, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
  • Psychology, acting, or writing skills that support unconventional adversarial thinking.
Work Terms
  • Remote, hourly engagement.
  • All work is text-based.
  • Higher-sensitivity projects are optional. Topics are communicated before exposure, with clear guidelines and wellness resources available.
Compensation
  • $48 to $62 per hour.
Impact
  • Identify vulnerabilities that automated testing may miss.
  • Expand evaluation coverage across more scenarios and help reduce unexpected production behavior.
  • Deliver reproducible safety artifacts that help make AI systems more robust, safe, and trustworthy.

Related Jobs

More like this