RedTeam Without Harm
An evaluation suite that stress-tests generative models against attempts to produce child-exploitation content using only synthetic, legal prompts — no real CSAM ever touches the system.
AI red-team leads and trust-and-safety engineers at model labs
- Library of thousands of synthetic attempt patterns that mimic known evasion tactics
- Refusal-rate and partial-compliance benchmarking across model versions
- Adversarial-prompt catalog updated monthly from disclosed incidents
- Compliance-ready report exports for EU AI Act and internal audit
The paper's second core contribution is a framework for red-teaming under ethical constraints; vendors are scrambling to translate the framework into shippable tooling before regulators mandate it.
Microsoft red-teamed 100 generative products in 2025, AI red-team engineers command $115K–$195K+, and IWF's 2025 report flags rising AI-generated CSAM — clear active investment in this domain.3 takeaways from red teaming 100 generative AI products ↗AI Red Team Engineer Job Description, Salary & Career Outlook ↗AI-Generated CSAM Trends | 2025 IWF Data & Insights Report ↗
No vendor appears to productize CSAM-generation red-teaming specifically; Thorn/Safer and Giskard cover adjacent detection and general LLM safety but not a shippable CSAM-stress-test suite — yet Thorn's domain authority and Microsoft's internal team are formidable indirect competitors.CSAM Detection from Experts in Child Safety Technology | Safer by Thorn ↗AI Safety at DEFCON 31: Red Teaming for Large Language Models (LLMs) ↗
Enterprise buyers (frontier labs, big-tech T&S teams) pay premium salaries for red-teamers and would license an eval suite, but the addressable market is narrow (handful of frontier model labs) and procurement is gated by safety/compliance budgets rather than product-led growth.AI Red Team Engineer - The $400K Career Path ↗3 takeaways from red teaming 100 generative AI products ↗
AI-generated CSAM is a structurally worsening threat flagged by IWF, Public Safety Canada, and Thorn; regulators (EU AI Act, US state laws) are pushing mandates, so demand for this category of tooling is durable and likely to compound.
Building the suite is non-trivial: needs curated synthetic prompt corpora, frontier-model API access, ethical review board, and deep safety expertise — Thorn and Microsoft have multi-year head starts in data, talent and trust.