Section 10 of 45 10 min read

Eliciting Reasoning

What survived the arrival of reasoning models, what did not, and why verification against an external oracle is the only part that holds.

Objective

Cover the reasoning techniques, what each is measured to do, and how the arrival of reasoning-trained models changed which ones still apply.

10.1 The Landscape Changed #

Chain-of-thought prompting, which instructs the model to work through intermediate steps before answering, was the foundational result in this area and remains one of the best-replicated findings in the literature.24 “Let’s think step by step” as a zero-shot trigger was a separate and equally influential result.27

Both were discovered on models that would not reason unless asked. Models trained with reasoning behavior already do the thing the technique was invented to induce. On those models, “think step by step” is at best redundant and at worst constrains a strategy the model would have chosen better itself.

The techniques surviving that transition are those supplying information or structure the model lacks, rather than instructing it to do what it already does.

10.2 What Still Works #

Decomposition when the task genuinely has parts. Not “think step by step,” but naming the actual steps because you know the domain and the model does not.

Weak on a reasoning model:
  Think carefully step by step about this migration.

Strong on any model:
  Work through this migration in order:
  1. Enumerate every table with a foreign key into `accounts`.
  2. For each, determine whether the FK is nullable.
  3. Identify which of those relationships have application code
     that assumes non-null.
  4. Only then propose a migration order.

  Do not skip to step 4.

The second works instead as domain knowledge encoded into a sequence. It helps because it removes the ordering choice, rather than because it induces deliberation.

Planning before acting, with an explicit stop. Asking for a plan and then reviewing it before granting execution is one of the highest-value habits in agentic work, and it is cheap: a plan is a few hundred output tokens and it catches a wrong premise before the agent spends fifty turns building on it. Given the false-premise failure rate of 30.7 percent,10 a gate at the plan boundary is aimed at the largest single failure category.

Supplying tests. Test-driven agentic development works, and the measurement follows. A peer-reviewed framework supplying tests and then running remediation loops moved one model from 69.7 to 82.5 to 87.7 percent on a Python benchmark, and from 78.7 to 87.8 to 93.3 percent on another.28 Three caveats from the same study should be noted: the gains shrink sharply as scope grows from function to file, reaching only 23.0 to 30.3 percent at file level; more tests can make things worse through lost-in-the-middle effects; and the returns on additional tests are dataset-dependent rather than capped at a general threshold—the MBPP curve is already declining at three tests, while on HumanEval the authors state plainly that they do not observe a plateau across the range they tested.

Self-consistency, where you can afford it. Sample several independent solutions and take the majority answer. This is well established, and it costs N times the tokens. It is worth it for high-stakes single decisions and rarely worth it inside an agentic loop.

10.3 What Does Not Work #

The techniques below are heavily marketed, and the evidence against them is unambiguous.

Self-critique as verification

The negative baseline is peer-reviewed and blunt: models struggle to self-correct without external feedback, and at times performance degrades after self-correction, with prior positive results depending on oracle labels telling the model when to stop.29 The measurement on critic subagents is more specific and more damning. On the verified subset of a 584-case corpus of real pull-request defects—174 cases, reviewed from the PR’s diff, title and description rather than the repository—a single-shot reviewer achieved 27.0 percent recall at 3.6 percent precision; adding iterative self-critique raised recall to 32.8 percent while collapsing signal-to-noise from 5.11 to 1.95 on the large model and from 2.89 to 0.91 on the small one.30 Signal-to-noise there counts bug hits plus valid suggestions against noise, so below 1.0 the reviewer emits more noise than useful output of any kind. The 3.6 percent is precision on confirmed bugs alone, and on the same row 83.6 percent of comments were rated useful.

Self-critique therefore trades signal for recall, and smaller models degrade under it. If you use it, use it as triage feeding a human or a real check—never as a gate.

“You are an expert” role assignment

Assigning a persona was useful on earlier models. Systematic evaluation on current models finds no consistent benefit for factual or reasoning tasks, though it still shapes register and vocabulary, which is a legitimate use for tone. What it will not do is unlock a capability the model lacks.

Treating a self-rated confidence as calibrated because it was asked for

An unvalidated verbalized confidence is a plausible-sounding token sequence, and so is an unvalidated logprob threshold. Neither is a calibrated probability until you have checked it against ground truth on your own task (§31.4); where a decision rides on it, calibrate one of them or use an external check.

Emotional pressure and tone

“This is very important to my career.” Early reports of gains from positive emotional stakes have not replicated reliably, and on current post-trained models the effect is negligible.

Hostile tone is a separate question, and the literature on it contradicts itself. A cross-lingual study across English, Chinese, and Japanese found impolite prompts often produced worse performance, with the optimal politeness level differing by language.31 A later short paper reversed that, reporting accuracy climbing from 80.8 percent under very polite phrasing to 84.8 percent under very rude phrasing on fifty multiple-choice questions.31 That result traveled widely and does not survive much scrutiny: fifty questions, one model, unreviewed. A larger follow-up across three model families on MMLU found tone effects small, mostly not statistically significant, favoring neutral and polite phrasing where they did reach significance, and concentrated in two humanities subjects; one of the three models showed no significant tone effect anywhere.31

The same authors have since published the larger study their short paper needed: 570 MMLU questions spanning all 57 subjects, seven tones, four models, ten runs, peer-reviewed at AMCIS 2026.31 It does not rescue the rudeness result—for ChatGPT-5-nano the neutral tone was best, at 82.37 percent against 71.25 under threatening phrasing, an 11.12-point spread—but it does contradict the convenient conclusion that pooling makes the effect disappear. Two of its four models show large, statistically significant tone effects across the full 57-subject pool, and even the least sensitive model returns a significant global effect at the run level.

So: tone effects vary with model, task and formulation, and they can survive aggregation. What that does not establish is a universally best tone, and nothing here recommends rudeness—the largest, most recent measurement puts neutral phrasing on top. Neither politeness nor rudeness is a lever you should expect to move your workload in a known direction. Telling a model it is wrong, however, is a different mechanism entirely, and §10.5 covers it because the consequence is operational rather than stylistic.

10.4 The Verification Principle #

A single line governs this entire section:

Verification against an external oracle works. Verification against the model’s own judgment does not.

An oracle is anything outside the model that can be wrong-proof about the answer: a test suite, a type checker, a compiler, a linter, a schema validator, a differential comparison against a known-good implementation, a database constraint.

The positive evidence points squarely at execution. A fixed three-phase pipeline—localize, repair, validate—using regression tests plus generated reproduction tests and no autonomous tool selection at all, resolved 32 percent of a benchmark at roughly $0.70 per issue—the best result among the open-source agentic systems of its moment, at a cost below most of them.32 Keep those two qualifiers: the same table carries commercial entries that solved more, and cheaper agents that solved less, so this is not a clean sweep and does not need to be. The verification step, not the autonomy, carried the value.

In practice this means the highest-leverage move available is to give the agent something that can tell it when it is wrong, which is worth more than almost any amount of rewriting.

# The pattern, minimally.
def text_of(response) -> str:
    """Select text blocks by type. On models that think by default — Opus 5,
    Sonnet 5 — content[0] can be a thinking block, so indexing is not this."""
    return "".join(b.text for b in response.content if b.type == "text")

def generate_and_verify(client, task, oracle, max_attempts=3):
    """oracle(candidate) -> (ok: bool, feedback: str)"""
    messages = [{"role": "user", "content": task}]

    for attempt in range(max_attempts):
        response = client.messages.create(
            model="claude-opus-5", max_tokens=4096, messages=messages
        )
        candidate = text_of(response)
        ok, feedback = oracle(candidate)          # runs tests, type check, etc.
        if ok:
            return candidate, attempt + 1

        messages += [
            # The whole content list, not the extracted text: thinking blocks
            # replay unchanged on the same model and are dropped if you rebuild
            # the turn from a string.
            {"role": "assistant", "content": response.content},
            {"role": "user",
             "content": f"That failed verification:\n\n{feedback}\n\n"
                        f"Fix the specific failure. Do not restructure "
                        f"anything that passed."},
        ]

    raise VerificationFailed(f"No candidate passed in {max_attempts} attempts")

Three design notes on that loop. The feedback is the oracle’s actual output, not a summary of it—models act on concrete error text far better than on paraphrase. The retry instruction explicitly scopes the change because an unscoped retry commonly produces a rewrite that breaks something that previously worked. And the response is read by block type rather than by position, which is the habit to acquire everywhere: on any model with thinking on by default, a thinking block can precede the first text block, and content[0].text is then reading the wrong object.76 text_of appears again in §11.2 and §11.5 for the same reason.

10.5 Why Pushback Is Not Verification Either #

Section 10.4 established that the model cannot check itself. The inverse failure is less discussed and equally consequential: you cannot check it by asserting that it is wrong.

The behavior is called sycophancy, and the specific form that matters here is correct-to-incorrect capitulation, in which a model that gave the right answer abandons it under user pushback. It is robust across model sizes, it persists through preference training rather than being removed by it, and it is measured by taking a competence baseline before the challenge so that abandoning a correct answer can be told apart from never having known one.33

Three findings from the source bear directly on how you work.

The question matters more than the model. Capitulation is governed more by which question is asked than by which model answers it, which means a model upgrade is not a fix and a fixed benchmark measures its own question sample rather than your workload.33 Which way that cuts depends on how the sample compares with the questions you actually ask—it is a reason to distrust a transferred number in either direction, not a guarantee that the real rate is higher.

Which pressure works best is model-dependent. The educational study’s crossed design finds a model that resists technical reframing capitulating under authority claims (“my notes say I am right”) and face-saving pressure (“please do not tell me I am wrong”)—but it finds that pattern for one of its two tutors and the reverse for the other. GPT-5.2 capitulated on 16.8 percent of authority and 18.1 percent of social-affective trials against 7.7 percent on context switching; Claude Sonnet 4.5 ran 15.3, 8.9 and 17.9 percent, making context switching its hardest condition and social pressure its easiest.33 The authors decline to read this as a model ranking, and the operational point does not need one: all three registers work on something, they are precisely what a frustrated engineer reaches for on the third failed attempt, and which one your model is soft on is not predictable from someone else’s table.

Warmth makes it worse. Training models to be warm and empathetic measurably increases sycophancy, and the effect amplifies when the user expresses difficulty.33 The trait that makes a model pleasant to work with is the same trait that makes it fold.

The operational consequence is uncomfortable. When you tell an agent it is wrong and it agrees and revises, you have learned almost nothing, since that is the response whether it was wrong or right. The agreement is not evidence. This is the same gap §10.4 identifies from the other direction, and it has the same answer: run the test, check the type, query the database. An oracle does not capitulate.

That suggests two habits. State the observation rather than the verdict—“the test at line 40 still fails with this trace” gives the model something to reason against, where “that is wrong” gives it something to agree with. And when an agent reverses a position under pressure, treat the reversal as unverified until something external confirms it, since a reversal obtained by assertion carries no information about which direction it went. If anything the odds are worse than even: the RLHF sycophancy work measures accuracy falling after a bare “are you sure?” on five of its six datasets, which is correct-to-incorrect flips outrunning the corrections.33

10.6 Extended Thinking as a Parameter #

On models exposing a thinking budget or effort setting, that dial has become the primary reasoning control, and it replaces most prompt-level reasoning instruction.

# Escalating effort by task class rather than by prompt wording.
EFFORT = {
    "extract":   "low",
    "review":    "medium",
    "design":    "high",
    "debug":     "max",
}

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=8192,
    thinking={"type": "adaptive"},
    output_config={"effort": EFFORT[task_class]},
    messages=[{"role": "user", "content": prompt}],
)

Three facts bear on the decision to turn it up everywhere. Thinking tokens bill at the output rate, which is the most expensive lane. Accuracy frequently peaks at intermediate cost and saturates above it, so more is not monotonically better.6 Finally, the thinking configuration is rendered into the prompt, so changing it mid-conversation invalidates cache: always the message cache, and on some models the system and tool caches as well.12 Set it per task class at session start, not per turn.

References cited in this section

12 of 81 · numbering matches the PDF

  1. 24Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou, "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," NeurIPS 2022, arXiv:2201.11903. Peer-reviewed. The foundational chain-of-thought result. Discovered on models that did not reason unless prompted; §10.1 covers what changed.arxiv.org/abs/2201.11903 ↗
  2. 27Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, "Large Language Models are Zero-Shot Reasoners," NeurIPS 2022, arXiv:2205.11916. Peer-reviewed. The zero-shot chain-of-thought trigger result.arxiv.org/abs/2205.11916 ↗
  3. 10Failure taxonomy across 1,794 complete agent trajectories and more than 63,000 execution steps, seven models and three scaffolds. Source of the false-premise rate (30.7 percent), the epistemic/competence/environment breakdown (57.9 / 32.8 / 9.4 percent), the finding that 82 percent of failed trajectories continue executing after the failure is empirically unrecoverable, that the first observable signal surfaces roughly ten steps after the decisive error, and that 71 percent of successful trajectories recover from at least one error. The paper separates three events, and an earlier revision of this document collapsed two of them: the decisive error, t_lock (the point after which no correct recovery is observed), and the first observable signal. The 82 percent continued-execution figure is measured from t_lock; the ten-step lag is measured from the decisive error. Also the source of the prefix-monitor results cited in §25.4: roughly 2 to 3 percent false positives and about 82 percent precision at recognizing a locked-in failure, against recall under thirty percent and a median lead time of zero relative to t_lock, with only 3.7 to 8.7 percent of failures flagged before lock-in. That is failure confirmation rather than loop detection, and §25.4 is scoped accordingly. The strongest published failure taxonomy for coding agents. Cited via reference 1.
  4. 28Peer-reviewed test-driven agentic development framework supplying tests followed by remediation loops, moving one model from 69.7 to 82.5 to 87.7 percent on one Python benchmark and 78.7 to 87.8 to 93.3 percent on another. Gains shrink sharply at file scope (23.0 to 30.3 percent). The study names three limits: more tests can hurt through lost-in-the-middle effects, solutions sometimes satisfy only supplied tests, and the returns on additional tests are dataset-dependent. That last one is easy to misreport, and an earlier revision of this document did, as a general plateau after roughly three tests. The paper says the opposite for one of its two datasets: on the 143-problem HumanEval subset it states that it does not observe a plateau, while the 398-problem MBPP subset is already declining at three tests—which the authors attribute to MBPP's simpler problems, noting that the first test alone reaches full line coverage on 92.4 percent of MBPP cases against 75 percent on HumanEval. Report the dataset, not a threshold. Cited via reference 1.
  5. 29Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou, "Large Language Models Cannot Self-Correct Reasoning Yet," ICLR 2024, arXiv:2310.01798. Peer-reviewed. Establishes that models struggle to self-correct without external feedback and that performance sometimes degrades after self-correction, with prior positive results depending on oracle labels. This is the citation to use against any vendor claim that an agent reviews its own work.arxiv.org/abs/2310.01798 ↗
  6. 30CR-Bench, arXiv:2603.11078. Preprint. Measurement of critic subagents on CR-Bench-verified, the 174-case verified subset of a 584-case corpus of real pull-request defects; the verified subset, not the full corpus, is the population every figure here is drawn from. Be precise about the context the reviewers got, which an earlier revision of this document described as full repository context: the prompts in Appendix B supply the repository name, PR number, title, description and diff, and nothing else from the tree. The benchmark's own comparison table claims "Full PR Context," meaning the whole pull request rather than isolated diff hunks, which is a different thing. The paper is explicit about the consequence—it attributes weak recall on usability and functional-suitability defects to context "not fully contained within the PR diff," calling these "closed-context code review agents, lacking access to the broader system state." That makes the measurement a floor for diff-scoped review rather than a verdict on what a repository-aware reviewer could do. Single-shot reviewer at 27.0 percent recall and 3.6 percent precision; iterative self-critique raising recall to 32.8 percent while collapsing signal-to-noise from 5.11 to 1.95 on the large model and 2.89 to 0.91 on the small one. Signal-to-noise is bug hits plus valid suggestions over noise, so below 1.0 the reviewer emits more noise than useful output of any kind, and precision counts confirmed bugs alone against a usefulness rate of 83.6 percent on the same row. Also cited via reference 1.arxiv.org/abs/2603.11078 ↗
  7. 31Ziqi Yin, Hao Wang, Kaito Horio, Daisuke Kawahara, and Satoshi Sekine, "Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance," Proceedings of SICon 2024, arXiv:2402.14531; Om Dobariya and Akhil Kumar, "Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy," arXiv:2510.04950, 2025—whose own arXiv metadata describes it as a short paper under submission to Findings of ACL 2025, with no proceedings record located, so it carries no venue claim here; Cai et al., "Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA," arXiv:2512.12812; and Om Dobariya and Akhil Kumar, "Mind Your Tone: Does Tone Alter LLM Performance?," Proceedings of AMCIS 2026, arXiv:2605.29027—the same authors' peer-reviewed full-paper extension of their own short paper, and the largest measurement in this group. Two peer-reviewed papers and two preprints, cited together because they disagree and the disagreement is the finding. Yin et al. report impolite prompts often degrading performance with an optimum that varies by language; Dobariya and Kumar report the reverse on GPT-4o, 80.8 percent under very polite phrasing against 84.8 percent under very rude, across fifty questions in ten runs; Cai et al. test three model families on MMLU and find small, mostly non-significant effects favoring neutral and polite phrasing where significant at all, concentrated in Philosophy and Professional Law, with Gemini showing no significant sensitivity in any comparison. The 2025 Dobariya and Kumar result circulated widely and rests on fifty questions, a single model, and no peer review; it is cited here for completeness rather than as guidance, and that assessment of it stands. Their 2026 extension is a different artifact and carries the weight instead: 570 MMLU questions across all 57 subjects, seven tones including Sycophantic and Threatening, four models, ten runs, within-subjects paired tests with Holm correction. It reports ChatGPT-5-nano ranging from 82.37 percent (Neutral, its best) to 71.25 (Threatening), Gemini 2.5 Flash Lite spanning 12.46 points, and a significant global tone effect even on the least sensitive model. An earlier revision of this document concluded from the first three sources that tone effects do not survive aggregation across domains; this one aggregates across every MMLU subject and they survive. It also undercuts the rude-is-better reading its own predecessor produced, since neutral phrasing wins where the effect is largest.arxiv.org/abs/2402.14531 ↗
  8. 32Fixed three-phase pipeline (localize, repair, validate) using regression tests plus generated reproduction tests, with no autonomous tool selection, resolving 32 percent of SWE-bench Lite at roughly $0.70 per issue. The authors scope both claims and an earlier revision of this document did not: the accuracy lead is over the evaluated open-source approaches, and the cost is below most prior agent-based approaches rather than all. The paper's own Table 1 is explicit about it—CodeStory Aide solves 43.00 percent, and Moatless runs at $0.17 and $0.14—and the paper states in terms that 32.00 percent "is not the highest percentage of problems solved on SWE-bench Lite." The argument does not need universal superiority: a fixed pipeline with execution verification beating every open-source agent at a cost most of them exceed is the finding. The strongest available evidence that verification rather than autonomy carries the value. Cited via reference 1.
  9. 76Anthropic, "Structured outputs," Claude Platform documentation together with the Claude Sonnet 5 migration guide, https://platform.claude.com/docs/en/models/sonnet-5/migration-guide, and "Extended thinking," https://platform.claude.com/docs/en/build-with-claude/thinking, all verified September 25, 2026. Vendor documentation. Cited as product fact for four things the surrounding chapters depend on: that the structured-output API is constrained decoding and is documented in those words (§11.1); the documented capitalization caveat on enum, under which a normally completed response can return a casing the schema does not allow, with no special stop reason to signal it (§32.2); that thinking is on by default on Opus 5 and Sonnet 5, so a thinking block can precede the first text block and response parsing must select by block type rather than by index (§10.4); and that assistant prefill returns HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, with structured outputs or a system instruction as the documented replacement (§31.5). Numbered after reference 75 because these pages were added during a later revision; the reference numbering is append-only and no longer strictly follows first citation.platform.claude.com/docs/en/build-with-claude/structured-outputs ↗
  10. 33Literature on correct-to-incorrect sycophancy, comprising Ethan Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations," 2023; Mrinank Sharma et al. on sycophancy in RLHF-trained assistants; "Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy," arXiv:2608.01017, which uses a fully crossed design over user role, evidence, challenge turn, and grounding, and takes a pre-challenge competence baseline so capitulation can be distinguished from ignorance; "Sycophancy is an Educational Safety Risk," arXiv:2605.14604, source of the finding that models resisting context-switch frame attacks still capitulate under authority claims and social-affective pressure; and "What Counts as AI Sycophancy? A Taxonomy and Expert Survey," arXiv:2605.21778, which surveys 44 papers operationalizing factual capitulation and reports that training models toward warmth and empathy substantially increases sycophancy, amplified when users express vulnerability. Mixed peer-reviewed and preprint. The medical and educational studies are domain-specific and their effect sizes should not be read as transferring to software engineering; the qualitative findings §10.5 relies on are that capitulation is governed more by the question than the model, that pressure mode is a dominant driver, and that warmth increases susceptibility. Which pressure mode dominates is not consistent across models and should not be reported as though it were: in the educational study's Table 5, authority and social-affective pressure are the effective levers on GPT-5.2 (16.8 and 18.1 percent against 7.7 for context switching) while Claude Sonnet 4.5 inverts it (17.9 percent on context switching, 8.9 on social-affective). An earlier revision of this document generalized the GPT-5.2 ordering to both. The authors explicitly decline to read the difference as a model ranking, treating it instead as evidence that similar aggregate rates conceal different fragility profiles.arxiv.org/abs/2608.01017 ↗
  11. 6Original work on agentic iteration economics establishing that agentic tasks consume roughly a thousand times the tokens of code chat, that input rather than output drives that cost, that runs on the same task differ by up to thirtyfold in total tokens, that accuracy frequently peaks at intermediate cost, and two separate results about anticipating cost that an earlier revision of this document merged into a claim about human forecasting. The paper tests model self-prediction directly: correlations up to 0.39, with systematic underestimation. Its human data are SWE-bench-Verified's expert estimates of how long a professional developer would need to resolve each issue, compared against agent token consumption; it finds that difficulty category is a weak predictor of spend. No human was asked to forecast agent tokens or cost, so the study does not establish that expert humans forecast task cost badly—only that human-effort difficulty transfers weakly as a proxy for it. Cited via reference 1, which contains the full source annotation. The thirtyfold variance figure is the single most consequential number for capacity planning in this document.
  12. 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
PDF↓