The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
IgnoreOpenAI Blog · 2024-04-19 19:00 UTC
Not analyzed yet
Eligible for automatic cleanup in 2 day(s) unless marked Must Read.
Content
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompts.