PoisonGarden
A managed data-poisoning service for small content sites that quietly feeds scrapers nonsense or watermarked versions of your pages so your work is harder to train on without blocking real readers.
Independent journalists, photographers, and niche bloggers who want to discourage AI scraping of their work
- Automatic invisible watermarking of text and images served to non-human visitors
- Optional trap pages full of plausible-but-wrong content that poison datasets without affecting real readers
- Monthly 'who scraped me' report showing scraper origins and payloads
- Per-page opt-in so you can protect your best work while leaving casual posts open
With residential proxy scrapers now overwhelming sites like LWN, individual creators are looking for ways to protect their work without killing their search traffic; this turns the LWN article's 'data poisoning' idea into a turnkey product.
Real and growing concern — LWN scraper piece (146 pts/139 comments), MIT Tech Review and Computerworld coverage, Cloudflare blog on residential proxies — but mostly framed as DDoS/cost issue, not a paid 'poison' service specifically; target niche bloggers/photographers have low budget tolerance.Fighting the AI scraperbot scourge - lwn.net ↗'Data poisoning' anti-AI theft tools emerge — but are they ethical? ↗
Clear differentiation: Glaze/Nightshade (U Chicago) cover images only, are free+open-source, and require manual artist-side processing; Cloudflare/DataDome block rather than poison; no commercial managed 'serve-nonsense-to-scrapers' CDN for text sites exists.Nightshade: Protecting Copyright ↗Web Scraping Protection - DataDome ↗5 Tools to Protect Your Content from AI Scraping and Poisoning ↗
Free open-source alternatives (Glaze, Nightshade) set the price anchor at $0; target users (independent journalists, niche bloggers, photographers) are notoriously budget-tight; enterprise scraping-protection pricing (DataDome/Cloudflare Bot Mgmt) is far above what small sites will pay.Nightshade: Protecting Copyright ↗Web Scraping Protection - DataDome ↗
Scraping will continue as long as models need data, so this stays relevant
Non-trivial: requires accurate scraper-vs-human classification at edge (residential proxies make IP-based heuristics weak), continuous retraining against new scraper fingerprints, and content-rewriting infrastructure that doesn't degrade SEO/human UX; effectively re-implements a slice of Cloudflare Bot Management for a small audience.Using machine learning to detect bot attacks that leverage residential proxies - Cloudflare ↗How AI Poisoning Protects Digital Content From Data Scraping Bots ↗