Key details
- Role type
- Contract
- Compensation
- $48–$62/hr
- Work arrangement
- Remote
- Category
- technology
- Confirmed requirements
- 4
About this role
Role Overview
Help strengthen AI safety by probing conversational AI systems with adversarial inputs, identifying vulnerabilities, and producing actionable red-team data. This text-based remote role includes reviewing AI outputs involving sensitive subjects such as bias, misinformation, and harmful behavior. Participation in higher-sensitivity projects is optional, with topics disclosed in advance and supported by clear guidelines and wellness resources.
Key Responsibilities- Test conversational AI models and agents for jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
- Annotate failures, classify vulnerabilities, and flag systemic risks.
- Use taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that customers can use to improve their AI systems.
- Native fluency in both English and Norwegian is required.
- Prior red-teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to test systems rigorously using frameworks or benchmarks and communicate risks clearly to technical and non-technical audiences.
- Comfort adapting across projects and customers.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, and model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk analysis, including harassment or misinformation probing, abuse analysis, and conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
- Finding vulnerabilities that automated testing misses.
- Delivering reproducible artifacts that improve customer AI systems.
- Expanding evaluation coverage across more scenarios and reducing production surprises.
- Remote, hourly engagement.
- $48 to $62 per hour.
What to prepare before applying
- Remote work location.
- Native fluency in English and Norwegian.
- Prior red teaming experience, including AI adversarial work, cybersecurity, or socio-technical probing.
- Adversarial mindset and ability to push systems to breaking points.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.