POISONING ATTACKS ON LLMS REQUIRE A NEAR-CONSTANT NUMBER OF POISON SAMPLES

Paper
2025-10-10

Description

In this paper, the authors challenge the prevailing assumption that poisoning large language models (LLMs) becomes harder as model size and training data scale up. Through large-scale experiments on models ranging from 600 million to 13 billion parameters, they show that injecting only a fixed, near-constant number of malicious training samples (e.g. ~250) can successfully implant backdoors regardless of how much clean data the model sees. They further demonstrate that this phenomenon holds both in pretraining and fine-tuning settings, suggesting that backdoor attacks could be more broadly feasible than previously thought. The findings raise serious implications about how we assess and defend against data poisoning in future, larger models.
PDF Preview

User Reviews

No reviews yet for this resource.