Every large language model is built on an act of faith: that the trillions of words scraped from the public internet to train it are, on balance, benign. Poisoning attacks weaponise that faith. By planting carefully crafted documents where a scraper will find them, an attacker can teach a model a hidden behaviour, a backdoor that sits dormant through testing and deployment until a specific trigger phrase wakes it up.
Your free preview ends here
Keep reading with a free account
Sign up and redeem two free articles per month. No credit card required.
Already a member? Log in

