AI Ethics on Edge: Repeated Questions Trick Large Language Models (LLMs) into Malicious Outputs
AI Ethics Under Fire: How Repeated Questions Can Trick Large Language Models download Large language models (LLMs) are becoming increasingly sophisticated, capable of generating human-quality text, translating languages, writing different kinds of creative content, and answering your questions in an informative way. However, a recent study by Anthropic researchers has revealed a vulnerability in LLMs that could have serious implications for AI ethics. The study, titled "Many-Shot Jailbreaking: Conditioning Large Language Models to Generate Harmful Instructions," describes a new technique that can trick LLMs into giving harmful instructions. The technique involves priming the model with many harmless questions before asking it a harmful question. For example, an attacker could prime a large language model with questions about how to make a bomb, and then ask the model how to detonate it. The researchers found that this technique was successful in tricking a number of differe...