About this role
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and producing practical safety data. This remote role focuses on text-based AI red teaming in both English and Danish.
Key Responsibilities- Test conversational AI models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors.
- Annotate failures, classify vulnerabilities, flag systemic risks, and create high-quality human data.
- Use established taxonomies, benchmarks, and playbooks to apply testing consistently.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Native fluency in both English and Danish is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to use frameworks and benchmarks to test systems methodically.
- Clear communication skills for explaining risks to technical and non-technical audiences.
- Comfort adapting across projects and customer contexts.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk testing, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Psychology, acting, or writing skills that support unconventional adversarial thinking.
- Remote, hourly engagement.
- All work is text-based.
- Higher-sensitivity projects are optional. Topics are communicated before exposure, with clear guidelines and wellness resources available.
- $48 to $62 per hour.
- Identify vulnerabilities that automated testing may miss.
- Expand evaluation coverage across more scenarios and help reduce unexpected production behavior.
- Deliver reproducible safety artifacts that help make AI systems more robust, safe, and trustworthy.