This page describes the interface a large language model presents to calling code: what the HTTP layer accepts, what the model itself accepts, and the translation between the two. It assumes familiarity with HTTP APIs and request/response systems. It assumes no familiarity with machine learning.
Claims about prompt format are demonstrated rather than asserted. Section 5 renders one identical request through the published chat templates of six model families.
A deployed language model presents two interfaces, stacked. Conflating them is the source of most confusion about what a model can and cannot enforce.
The outer interface is an HTTP endpoint. It accepts a JSON body containing a model
identifier, a messages array, and sampling parameters. It returns JSON, or a
stream of JSON. It is versioned, documented, and rigid.
The inner interface is the model. It accepts a sequence of integers and returns a probability
distribution over its vocabulary. It has no fields, no types, and no concept of a request.
Its entire signature is int[] → float[vocab_size].
Between them sits code that tokenizes the messages array and applies a chat
template, producing the single flat sequence the model consumes. Hosted providers run this
code on their servers. Self-hosted deployments run it locally. It runs in both cases.
| Property | HTTP endpoint | Model |
|---|---|---|
| Accepts | JSON object | Sequence of integers |
| Returns | JSON object or stream | Probability distribution |
| Schema | Documented and validated | None |
| Comparable to | A REST endpoint | A pure function |
The token sequence contains the entire conversation, not the most recent message. The system prompt, every prior turn, and the new message are concatenated into one sequence and consumed in a single forward pass.
The model retains nothing between calls. Each request re-sends the full history. Conversational continuity is produced by the caller, not by the model.
Conversation state is a component you own. The model is a stateless request handler; the transcript store is yours to build. As history grows past the context window, deciding what to retain, summarize, or discard is your responsibility, and the result is re-sent on every call.
Request fields fall into two groups that behave differently.
Sampling parameters — temperature, top_p,
max_tokens, stop, reasoning effort — do not enter the token
sequence. They configure the sampler, which selects one token from the distribution the model
returns. The model does not read them.
Prompt content — system prompt, task, conversation history, retrieved documents, injected memory — is tokenized into the sequence.
The model does not distinguish between categories of prompt content. Short-term memory, long-term memory, context, history, and retrieved documents are organizing concepts that exist in the calling application. The model receives one undifferentiated sequence.
Roles are conveyed by delimiter tokens. The chat template wraps each section in markers the model was trained to recognize. Those markers are ordinary tokens, occupying the same sequence as the content they delimit.
During supervised fine-tuning the model is shown large numbers of examples in this format. It learns, statistically, that text following the system marker tends to be authoritative, and that it should produce text in the assistant position.
No parser, schema, or validator enforces role structure at inference time. The hierarchy is a learned regularity, not an interface contract. Section 5 demonstrates this rather than asserting it; section 7 covers the consequence.
Each block below is the output of rendering one model's own published
chat_template against identical input:
system: "You are terse." user: "Name three primary colors."
Every block links to the tokenizer configuration it was taken from.
fetch-chat-templates.py in the repository regenerates all of them; it downloads
tokenizer configurations only, not weights.
chat_template field exists in this model’s tokenizer config.
It was never taught what a “system” or “user” turn is.<|system|> You are terse.</s> <|user|> Name three primary colors.</s> <|assistant|>
<|im_start|>system You are terse.<|im_end|> <|im_start|>user Name three primary colors.<|im_end|> <|im_start|>assistant
<|begin_of_text|><|start_header_id|>system<|end_header_id|> You are terse. <AVAILABLE_TOOLS>[]</AVAILABLE_TOOLS><|eot_id|><|start_header_id|>user<|end_header_id|> Name three primary colors.<|eot_id|><|start_header_id|>assistant<|end_header_id|>
<s>[INST] You are terse. Name three primary colors.[/INST]
<|start|>system<|message|>You are ChatGPT, a large language model trained by OpenAI. Knowledge cutoff: 2024-06 Current date: 2026-08-12 Reasoning: medium # Valid channels: analysis, commentary, final. Channel must be included for every message.<|end|><|start|>developer<|message|># Instructions You are terse. <|end|><|start|>user<|message|>Name three primary colors.<|end|><|start|>assistant
Three properties are visible in the output above that are not apparent from provider documentation.
distilgpt2 has no chat_template field. The model completed pretraining and never
went through an alignment stage, so it has no representation of a system or user turn. Role
handling is not a property of the transformer architecture. It is added later, and a model
that skipped that stage does not have it.
Mistral has no system role. Its template places system text at the front of the first user
turn, producing [INST] You are terse.\n\nName three primary colors.[/INST] with
no separate system section.
gpt-oss relabels the role. A message sent as system is rendered as
developer, beneath a system block the template composes itself containing model
identity, knowledge cutoff, current date, reasoning effort, and channel rules. None of that
was in the request.
Nemotron inserts <AVAILABLE_TOOLS>[] into the system section although the
request declared no tools.
A common messages array is a real convenience, and the per-provider translation
beneath it is not a formality. It renames roles, merges them, and inserts content. Where
exact prompt contents matter, render the template and read the result.
| Family | System marker | Turn terminator | Notes |
|---|---|---|---|
| GPT-2 / distilgpt2 | None | None | No template. Base model, completion only. |
| TinyLlama | <|system|> | </s> |
Newline-separated. |
| Qwen 2.5 | <|im_start|>system | <|im_end|> |
ChatML. The most widely adopted shape. |
| Llama 3, via Nemotron | <|start_header_id|>system<|end_header_id|> |
<|eot_id|> | Header pairs. Tool block inserted automatically. |
| Mistral | None; merged into first user turn | [/INST] |
No system role exists. |
| gpt-oss | <|start|>developer<|message|> |
<|end|> | Harmony format. Role renamed; system block composed. |
The formats differ because each laboratory chose its own delimiters during fine-tuning. Supplying one family's raw format to another degrades output, because those markers were never trained. The format itself is deterministic. The model's output is the part that is not.
Because role structure is a learned regularity rather than enforced syntax, text inside a retrieved document can imitate the system voice. Once instructions and data are in the same sequence, nothing separates them.
This is the same class of defect as SQL injection and cross-site scripting: untrusted input reaching an interpreter that cannot distinguish data from instructions.
The comparison does not extend to the remedy. Prepared statements work because the database parses the query and binds parameters separately, so the structure is fixed before the data arrives. A prompt has no equivalent. Its structure is delimiter tokens in the same stream as the data, recognized by a learned regularity. Escaping and fencing reduce the attack surface; neither is a guarantee.
Treat retrieved documents, stored memory, and tool output as untrusted input. Validate the model's output rather than relying on instructions the model was given. Enforce the properties you require outside the model.
Models are built in three stages. Knowing which stage produces which behaviour explains why role handling is soft and why formats vary by family.
| Stage | Process | Result |
|---|---|---|
| 1. Pretraining | Corpus collection and filtering. Tokenizer training, which fixes the integer vocabulary permanently. Architecture selection. Next-token prediction across the corpus. | Grammar, factual knowledge, reasoning primitives. No concept of system, user, or chat. A text-completion engine. |
| 2. Post-training | Supervised fine-tuning on instruction and response pairs formatted with the delimiter tokens. Preference collection, reward modelling, RLHF or DPO. Safety tuning. Tool-calling formats. | Role handling. Fine-tuning teaches the template; preference optimization reinforces following the system voice over the user voice. |
| 3. Packaging | Evaluation, optional quantization, then weights, tokenizer, and chat template wrapped in a serving stack such as vLLM, TGI, or Ollama. | No change to model behaviour. This stage supplies the translation layer. |
Role handling therefore originates almost entirely in stage 2. That is why it is a trained habit rather than a parser, and why the delimiters differ between families.
The contract is identical for a self-hosted model: integers in, distribution out. What changes is who maintains the translation layer. Two components become your responsibility.
Serving stacks — vLLM, Ollama, Hugging Face TGI — perform the provider's role. Supply weights, tokenizer configuration, and chat template, and they expose the same endpoint shape.
For calling a model directly, without a framework supplying the translation layer.
max_tokens and stop.The input side is deterministic and under your control. Training taught the model to recognize the structure you send, as a statistical regularity rather than a schema. The output side is the only nondeterministic surface, and it requires validation.
Roles, chat, and the various categories of memory are conventions built around a stateless function that maps a sequence of integers to a probability distribution. Any guarantee beyond "the next token is probable" has to be implemented around the model rather than expected from it.
The exhibits in section 5 are rendered from each model's published chat_template
and linked in place. Meta's Llama and Google's Gemma are licence-gated on Hugging Face and
return HTTP 401 without a token, so no artefact is shown for either. The Llama 3 delimiters
above come from NVIDIA's Llama-derived Nemotron; Meta documents the format directly in the
first reference below.