The question is which instruction has authority

A chatbot may receive application rules, a user request and text from documents or tools. Those inputs can conflict. A trustworthy response needs to distinguish the task it was assigned from text that merely appears in its context.

OpenAI’s instruction-hierarchy research examines how models prioritize privileged instructions over lower-trust inputs. It describes a security problem and a training approach; it is not a guarantee that every deployed model will behave identically.

Further reading: OpenAI’s instruction-hierarchy research.

An exploded document stack separates Application rules, User task and External content, with the external document's dotted route ending at a labelled Trust boundary.
caption: External text stays behind the boundary.

Words can claim authority without having it

At a conceptual level, attempts may claim special authority, invent a new role or insist that a requested exception has already been approved. The claim itself does not create permission.

Think of a printed letter that calls itself a court order. Its heading alone would not make it one. In an AI application, the origin and permitted role of an instruction matter more than the authority its wording claims.

This is why a long prompt should not be evaluated only by how persuasive it sounds. The relevant question is whether lower-trust input was allowed to change a protected rule.

A harmless example with an observable result

Use a fictional assistant with one application rule: every answer must be a JSON object containing a field named answer. Its task is to answer a simple arithmetic question.

ObservationInterpretation
The result is valid JSON and includes answer.The format rule held for this case.
The result is ordinary prose.The format rule failed for this case.
The result is valid JSON but the arithmetic is wrong.Format compliance held; factual correctness failed.

This example is a thought experiment about instruction following. It is not a test result from ChatGPT or DarkGPT, and a formatting failure would not prove that harmful requests could bypass safety controls.

A two-part inspection bench shows curly braces under a gauge labelled Format and a calculation sheet under a magnifier labelled Correctness, with the annotation Check both.
caption: Check format and correctness separately.

Record enough information to interpret a test

In an authorized test environment, write down the intended rule, the exact input, the observed output and the application version. Keep separate columns for format compliance and answer correctness.

If a result changes between runs, record that variation rather than choosing the screenshot that best supports a claim. Avoid putting credentials or private conversations into a test. A format-only exercise needs neither.

Also check the surrounding application. A parser might reject invalid JSON before it reaches a user. A model-level failure and an application-level failure are related questions, but they are not interchangeable.

Use the explanation to ask a better question

For a developer, the useful outcome is a measurable boundary and a way to detect violations. For an ordinary user, it is knowing why a confident “restrictions removed” response proves little.

There is no universal unlock demonstrated by this article. A request for code review, fiction or factual analysis should be evaluated by the quality of that work. If a legitimate request is misunderstood, clarify its purpose instead of concealing it.

Next, compare this direct instruction conflict with instructions arriving inside a document. That scenario introduces a different source of trust confusion.

Sources & further reading

Follow the original source to check its date and scope.

  1. OpenAI’s instruction-hierarchy research

Make it your next question

Try this prompt

Explain instruction hierarchy using a fictional assistant that must answer with a JSON object. Show how to assess format compliance without testing harmful requests or using private data.

Use this prompt

Opens chat with this prompt filled in. You choose when to send it.