OpenAI and Codex
OpenAI and Codex: configuration surface, caching model, prompt cache keys, pricing, and the distinctive features.
Cover the Responses API surface and Codex configuration.
39.1 Configuration Surface #
| Mechanism | Location | Loads |
|---|---|---|
| Agent instructions | AGENTS.md (nested supported) | Every turn |
| Rules | Codex rules configuration | Per rule scope |
| Config | ~/.codex/config.toml and project config | Startup |
| Skills | Skills directory / plugin bundles | On demand |
| Subagents | Subagent configuration | On delegation |
| Hooks | Codex hooks | On event |
| MCP | Codex MCP configuration | Schemas every turn, or on discovery with tool search |
| Memories | Codex memories | Every turn |
Codex reads AGENTS.md natively, which makes it the reference implementation of the standard.71 Nesting works as specified—nearest file wins.
39.2 Caching Model #
OpenAI’s model changed materially with GPT-5.6, and the differences matter.49
GPT-5.6 and later:
- Both implicit and explicit breakpoints supported.
prompt_cache_options.modeselects;prompt_cache_breakpointmarks. - Up to four cache writes per request. Lookup considers the first 2 and the latest 50 explicit breakpoints and reuses the longest match; implicit mode also considers the boundaries at the ends of messages.
- Cache writes at 1.25×, reads at 0.1×.
- TTL:
30mis the only supported value and the default. - Minimum cacheable prefix: 1,024 visible input tokens.
- Explicit-only mode avoids writing changing suffixes, which is a genuine cost control. What is distinctive is the mode selector that turns implicit boundaries off; the underlying control—mark the stable prefix, leave the changing suffix unmarked—is the same one §22.3 teaches and the same one Anthropic’s per-block
cache_controlprovides (§38.2).
Earlier models:
- Implicit only, breakpoints at model-dependent intervals.
- No cache write charge; model-dependent read rate.
prompt_cache_retentionofin_memoryor24h. These are maxima rather than working lifetimes:in_memoryis typically 5–10 minutes against a one-hour ceiling,24htypically around 30 minutes against a 24-hour one. On models offering both, the default follows the organization’s data-retention policy—24hwithout Zero Data Retention,in_memorywith it. GPT-5.5 and 5.5 Pro support24honly.- Minimum cacheable prefix varies with the request settings;
cached_tokensrounds down to a multiple of 128.
Follow the migration checklist from OpenAI’s own documentation literally: keep stable prefixes, keep prompt_cache_key values, replace prompt_cache_retention with prompt_cache_options.ttl, confirm the prefix meets the new minimum, and add an explicit breakpoint after stable content if the default breakpoint would include changing material.49
39.3 Prompt Cache Keys #
This is the distinctive OpenAI mechanic, covered in §22.7. Cached states are machine-local and traffic above roughly 15 requests per minute can overflow. On models before GPT-5.6, prompt_cache_key influences routing so related requests reach the same cache; from GPT-5.6 the routing is automatic and the key is optional, marking off separate cache accounting instead.
const response = await client.responses.create({
model: "gpt-5.6-sol",
prompt_cache_key: "triage-v3:tenant-acme:shard-7",
prompt_cache_options: { mode: "explicit", ttl: "30m" },
input: [
{ role: "developer", content: [
{ type: "input_text", text: STABLE_INSTRUCTIONS,
prompt_cache_breakpoint: { mode: "explicit" } },
]},
{ role: "developer", content: `Request time: ${new Date().toISOString()}` },
{ role: "user", content: userMessage },
],
});
The order runs stable content with an explicit breakpoint, then the varying developer content, then the user message. Explicit-only mode means the varying suffix is processed at the uncached rate with no write charge, which is exactly right for content that will not be reused.
39.4 Pricing #
Verified September 24, 2026, standard tier, per million tokens:52
| Model | Input | Cached | Cache write | Output | Long-context input |
|---|---|---|---|---|---|
| gpt-6-astra | $10.00 | $1.00 | $12.50 | $50.00 | $20.00 |
| gpt-5.6-sol | $4.00 | $0.40 | $5.00 | $20.00 | $8.00 |
| gpt-5.6-terra | $2.00 | $0.20 | $2.50 | $12.00 | $4.00 |
| gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 |
Batch is 50 percent off. Data-residency endpoints add a 10 percent uplift for models released on or after March 5, 2026. The long-context column is the threshold cliff from §21.3. It begins above 272K input tokens, and crossing it applies 2× to input, cached reads and cache writes, and 1.5× to output, for the full request.
39.5 Distinctive Features #
context_management replaces earlier conversation content with a compacted representation server-side. It reduces input tokens and can reduce cache reuse simultaneously—OpenAI’s guidance is to compare total input cost before and after, since fewer tokens can still save money even when the hit rate falls.49
Discovered tools append at the end of context, which preserves the earlier reusable prefix. This is the cleanest solution to the §17.1 fixed-block problem on any platform.
Restrict which tools are callable while keeping the tools list byte-identical, so the cache survives. Removing a tool definition invalidates the cache, whereas restricting the allowlist leaves it intact.
On GPT-6 Astra, a configuration_update input item changes reasoning effort between responses while leaving request-level reasoning.effort unchanged, preserving the prefix for cache reuse.
References cited in this section
3 of 81 · numbering matches the PDF
- 71OpenAI, "AGENTS.md," Codex documentation Vendor documentation; cited as product fact for Codex's native AGENTS.md support and nesting behavior.developers.openai.com/codex/agent-configuration/agents-md ↗
- 49OpenAI, "Prompt caching," OpenAI API documentation verified September 8, 2026. Vendor documentation; cited as product fact for implicit and explicit breakpoint modes, the 1.25× write and 0.1× read multipliers on GPT-5.6 and later, minimum cacheable lengths, TTL and retention semantics, machine-local cache routing and the ~15 requests-per-minute overflow threshold, prompt_cache_key design guidance, the minimum-cacheable-length break-even formula, and the compaction interaction.developers.openai.com/api/docs/guides/prompt-caching ↗
- 52OpenAI, "Pricing," OpenAI API documentation verified September 8, 2026, with the per-model pages under https://developers.openai.com/api/docs/models/ verified September 24, 2026. Vendor documentation; source for all OpenAI per-model rates including the short-context and long-context tiers. The 272K-token threshold and the 2× input, 2× cache and 1.5× output multipliers are stated on the model pages rather than on the pricing page, which lists the two tiers without naming the boundary. One rate in §21.3's table is promotional rather than standing: the gpt-5.6-sol model page documents its $4.00/$20.00 as available at least through November 21, 2026, verified September 25, 2026, and describes it as a reduction against the prior generation. Re-check it before carrying it into a forecast.developers.openai.com/api/docs/pricing ↗