Section 33 of 45 5 min read

Post-Training, and Why Certain Phrasings Work

Three stages of post-training, what preference training explains, why format matters more than wording, and the status of magic words.

Objective

Explain where prompt sensitivity comes from, so that technique selection is grounded in how models were shaped rather than in accumulated superstition.

33.1 Three Stages #

A deployed model has passed through roughly three stages, and each leaves fingerprints on which prompting techniques work.

Pretraining

on a large corpus produces next-token prediction. A pretrained base model merely completes text. Ask it a question and it may generate more questions because that is what follows a question in many documents.

Supervised fine-tuning

on curated instruction-response pairs produces a model that responds to instructions. The format of those pairs—a system message, a user turn, an assistant turn—becomes the format the model expects, which is why the chat template is not cosmetic.

Preference optimization

— whether RLHF, DPO, or a variant — trains the model toward outputs humans rated higher. This is the origin of most behaviors practitioners describe as “personality,” and where several prompting effects originate.

33.2 What Preference Training Explains #

Several widely observed behaviors follow directly from this stage rather than from anything intrinsic to language modeling.

Structured output preference

Raters preferred organized responses, so models produce headers and bullets readily. This is also why asking for a specific structure works so reliably: you are selecting among behaviors the model was rewarded for.

Hedging and length bias

Longer, more qualified answers were rated higher on average. This produces verbose responses to simple questions, and it is why explicit brevity instructions are necessary rather than merely helpful.

Sycophancy

Models tend toward agreement because agreement was rated well. This has a direct consequence for verification, and it compounds the §10.4 finding from both directions: a model asked to check its own work is operating under a trained tendency to validate, and a model told it is wrong is operating under the same tendency (§10.5).

Instruction-hierarchy sensitivity

Post-training establishes that system content outranks user content. This is a trained tendency, not an architectural guarantee, which is the entire premise of Section 36.

33.3 Why Format Matters More Than Wording #

The models were trained against specific templates. Matching the template shape puts the model in the distribution it was optimized on.

This explains why several structural choices outperform their wording-level equivalents:

  • XML-tagged sections work well because tagged structure appears throughout training data.
  • Markdown headers work well for the same reason.
  • Numbered steps work because instruction-following data is full of numbered steps.
  • A “system prompt in the user message” underperforms a real system message because it is off-template.

It also explains why elaborate wording tweaks are usually low-yield. You are moving within a broad basin, whereas format changes move between basins.

The size of that second effect is measured. Holding content constant and varying only the template across plain text, Markdown, JSON, and YAML shifted GPT-3.5-turbo by roughly 18 percent relative on code translation—the spread its published tables display, against the “up to 40 percent” its abstract claims and does not reconcile (reference 18)—and the sensitivity was statistically significant on nearly every dataset tested; larger models proved more robust without becoming insensitive.18 Work using a broader search over semantically equivalent formats reports swings reaching 76 percentage points on some open-weight models.19 No rewording exercise in my experience has produced a comparable movement, which is the practical case for the ordering of Part II.

The difference is visible in a pair. These are two versions of the same request, written the way each would actually arrive:

A. I need you to look at the config below and let me know if there's
   anything that stands out as being potentially problematic from a
   security standpoint, if you could.

   listen_addr = 0.0.0.0:8080
   tls = false
   admin_token = "changeme"
B. <task>
   Review the configuration for security defects. Report only defects
   that are exploitable as written. For each: the setting, the risk,
   and the corrected value. If there are none, say so.
   </task>

   <config>
   listen_addr = 0.0.0.0:8080
   tls = false
   admin_token = "changeme"
   </config>

B is not more polite, more emphatic, or more knowledgeable about the domain. It is closer to the shape of the instruction-following data the model was tuned on: a delimited task, a delimited input, an explicit output contract, and a stated exit condition.

Be clear about what that pair does and does not isolate, though. It is a combined rewrite rather than a format-only change: B also narrows scope to exploitable defects and adds an exit condition, so the task contract moved along with the shape. The controlled evidence for format alone is the measured work above, which held content constant and varied only the template.18,19 The pair is here to show what the two habits look like side by side, and the reason Part II is overwhelmingly about structure rather than phrasing is that measurement, not this illustration.

The corollary is a useful diagnostic. When a prompt underperforms and rewording has not helped, stop rewording. Check whether the instruction, the data, and the output contract are separated at all—in most underperforming prompts they are not.

33.4 What This Says About “Magic Words” #

Techniques that made a real difference on earlier models often no longer do, because the behavior they once induced has since been trained in.

TechniqueOriginStatus on current models
”Let’s think step by step”Zero-shot CoT27Redundant on reasoning models
”You are an expert…”Persona conditioningNo consistent benefit for accuracy
”This is important to my career”Emotional promptingNegligible; does not replicate
”Take a deep breath”Optimizer-discovered phraseModel-specific and largely obsolete
Explicit output format—Still works, reliably
Negative constraints—Still works, reliably
Supplying an oracle—Works and is the largest lever

The pattern here is clean. Techniques that induced a behavior have decayed as that behavior was trained in. Techniques that supply information the model does not have have not decayed at all, because no amount of training gives the model your constraints.

That is the durable frame for everything in Part II: information transfer ages well, while behavioral tricks decay as the behavior gets trained in.

33.5 Model-Specific Behavior #

Models from different labs, and successive generations from the same lab, respond differently to identical prompts. This is a consequence of different post-training data and objectives, and it is not going away.

The practical implications follow:

  • A prompt tuned on one model is not portable without re-evaluation.
  • A model version bump is a change requiring the eval suite to run (§34).
  • Vendor prompting guidance is model-specific, and the vendor knows what their post-training rewarded.
  • Cross-model portability requires prompting to the intersection: explicit structure, explicit constraints, explicit output contract, no reliance on any one model’s inferential habits.

Section 43 develops the portability case concretely.

References cited in this section

3 of 81 · numbering matches the PDF

  1. 18Jia He, Mukund Rungta, David Koleczek, Arshdeep Sekhon, Franklin S. Wang, and Sadid Hasan, "Does Prompt Formatting Have Any Impact on LLM Performance?," arXiv:2411.10541, November 2024. Preprint. Formats identical content into plain text, Markdown, JSON, and YAML and evaluates across natural language reasoning, code generation, translation, and named-entity tasks on four OpenAI models. Source of the abstract's claim that GPT-3.5-turbo varies by as much as 40 percent on a code translation task purely by template, and of the finding that format sensitivity is statistically significant (p < 0.05) on every dataset and model pairing tested except one. Carry the 40 percent as an unreconciled abstract claim, because the paper's own translation results do not display it: the widest GPT-3.5 spread in Tables 7–8 is Java-to-C# BLEU running from 66.46 under plaintext to 78.40 under JSON, which is 11.94 points or roughly 18 percent relative, and the paper's other headline number—accuracy up 42 percent for JSON over Markdown—is from MMLU international law rather than code translation. This is an internal inconsistency in the source rather than a misquotation, and no calculation reconciling the abstract with the tables appears in the paper. Where a defensible magnitude is needed, cite the table result; do not silently swap in the 42 percent figure, which is a different task and a different metric. Also the source for the observation that larger models are more robust without being insensitive, and that the best-performing template does not transfer reliably between models. The study varies whole-prompt template rather than inline emphasis markup, which is the limitation relevant to §8.4.arxiv.org/abs/2411.10541 ↗
  2. 19Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design: Or, How I Learned to Start Worrying About Prompt Formatting," ICLR 2024, arXiv:2310.11324. Peer-reviewed. Introduces FormatSpread, which searches over semantically equivalent prompt formats rather than testing a handful by hand, and reports accuracy differences of up to 76 percentage points attributable to formatting alone on tasks drawn from SuperNatural-Instructions. Cited for the magnitude of format sensitivity and for the methodological point that evaluating a model on a single format reports one sample from a wide distribution—the same argument §34.4 makes about running an eval once.arxiv.org/abs/2310.11324 ↗
  3. 27Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, "Large Language Models are Zero-Shot Reasoners," NeurIPS 2022, arXiv:2205.11916. Peer-reviewed. The zero-shot chain-of-thought trigger result.arxiv.org/abs/2205.11916 ↗
PDF↓