GPT-5.5’s Split Availability: A Documented API Review of the ChatGPT/Codex Retirement and the GPT-5.5 vs GPT-5.4 Trade-Off

GPT-5.5's Split Availability: A Documented API Review of the ChatGPT/Codex Retirement and the GPT-5.5 vs GPT-5.4 Trade-Off
GPT model trade-offs represented by balanced technical comparison forms
generative pre-trained transformer (GPT)A family of transformer-based language models trained to generate and analyze content. Open glossary entry model trade-offs represented by balanced technical comparison forms.

Evidence checkpoints

Documented point: As of the documentation date, GPT-5.5’s 14 October retirement is a ChatGPT, ChatGPT Work and Codex product-surface change, not an application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry retirement. The date is prospective as of 30 September 2026. Unaffected API availability does not mean identical account, region, endpoint or rate-tier access, and configurations selecting gpt-5.5 still need review. [official source 1]

Documented point: Listed standard GPT-5.5 text-token rates are exactly twice GPT-5.4: $5 versus $2.50 input, $0.50 versus $0.25 cached input, and $30 versus $15 output per million tokens. Unit price is not total workflow cost. Above 272,000 input tokens, both pages state higher long-context rates; regional processing has a stated uplift, and tools and retries also matter. [official source 1 official source 2]

Documented point: Both pages list a 1,050,000-token context window, 128,000-token maximum output, text and image input, text output, and broadly matching endpoint and tool support. Listed support does not establish equal quality, reliability, latency, tool behaviour, or account limits; recheck live docs at publication. [official source 1 official source 2]

Documented point: GPT-5.5 defaults to medium reasoning and GPT-5.4 to none; OpenAI warns that higher GPT-5.5 effort can add latency and cost and regress quality with conflicting instructions, weak stopping or open-ended tools. This claim rests on API documentation and vendor guidance, not on ChatGPT user-interface settings or a universal result. [official source 1 official source 2 official source 3]

Documented point: A defensible decision compares both on representative inputs for accepted quality, end-to-end latency, tokens, tool and retry behaviour, and total cost, retaining the lightest configuration that passes. This is a method, not a claim that either model wins every workload; fixed snapshots improve reproducibility. [official source 1 official source 2]

The dates describe different facts

30 September 2026 is the evidence cut-off

The 30 September date records when the current developer pages and Codex model notice were reviewed for this article. It is not a GPT-5.5 launch date, retirement date or guarantee that the same wording will remain current. Availability pages are operational documents and may change after an article is prepared. A production decision should therefore preserve the date of each documentation check and repeat that check at approval and rollout.

The evidence date matters because lifecycle statements are time-sensitive. On 30 September, saying “GPT-5.5 has retired from ChatGPT and Codex” would have been incorrect under the cited notice: 14 October was still in the future. Saying “GPT-5.5 is retiring from the API on 14 October” would also have been incorrect because OpenAI expressly excluded the API from that retirement notice.

A dated status statement should use four fields: the source checked, the access surface, the observation date and the effective date. For this case, the relevant record is: OpenAI Codex model documentation; ChatGPT, ChatGPT Work and Codex surfaces; checked 30 September 2026; retirement scheduled for 14 October 2026. The API belongs in a separate row marked unaffected by that notice, subject to live account and documentation verification.

14 October 2026 is the scheduled surface-retirement date

OpenAI’s wording makes 14 October a product-surface retirement date. “Product-surface retirement” means that a model ceases to be offered through specified user-facing or agent interfaces; it does not automatically mean that the corresponding API model identifier is deprecated or removed. Lifecycle notices must be read for both the named model and the named delivery channel.

OpenAI also notes plan-dependent replacements in its product documentation. That should not be converted into a guaranteed replacement for a particular user, workspace or Codex configuration. Administrators need to inspect the options actually exposed to their plans and workspaces, then verify permissions, defaults and workflow behaviour rather than assuming a replacement from general documentation.

The October deadline should therefore trigger a product-surface readiness check, not an unsupported declaration of API deprecation. Teams with both Codex and direct API workloads may update the Codex configuration while retaining a separately evaluated GPT-5.5 API deployment, because the lifecycle boundary is specific to each surface.

Documented status boundary as reviewed on 30 September 2026
Surface Documented status Operational interpretation
ChatGPT GPT-5.5 retirement scheduled for 14 October 2026 Review affected use before the scheduled date; do not describe retirement as already complete on 30 September.
ChatGPT Work GPT-5.5 retirement scheduled for 14 October 2026 Check plan, workspace configuration and the replacement actually available; do not guarantee plan-specific options.
Codex GPT-5.5 retirement scheduled for 14 October 2026 Identify ChatGPT-authenticated and other affected Codex configurations rather than assuming every API-backed implementation has the same lifecycle.
Direct OpenAI API Explicitly excluded from the cited retirement notice Continue a separate model-selection review, while confirming current account, region, endpoint and rate-tier access.

Why “ChatGPT”, “Codex” and “API” cannot be used interchangeably

ChatGPT is a product interface, ChatGPT Work is a distinct work surface, Codex is a coding product with its own model-selection and authentication arrangements, and the OpenAI API is the programmable service used by applications. A model name can appear across several of these surfaces without acquiring one shared lifecycle. Documentation that retires a model from one surface does not, by itself, establish retirement from another.

The distinction is especially important for Codex. A team may use Codex through ChatGPT authentication, use a configuration connected to an API key, or call the model directly from its own application. These routes should not be collapsed into a single “Codex/API” category. The retirement notice names product surfaces, while the model reference supplies API-specific information. Authentication and routing must be inspected in the actual configuration.

Recommended inventory method: record each workload’s user entry point, authentication mechanism, model selector, endpoint, model identifier, snapshot policy, tool dependencies and responsible owner. Then classify it as a ChatGPT, ChatGPT Work, Codex or direct API dependency. Where the route is unclear, treat that uncertainty as a migration blocker rather than inferring status from the model name alone.

  • User entry point: note whether work starts in ChatGPT, ChatGPT Work, Codex or an organisation-owned application.
  • Authentication: distinguish ChatGPT-authenticated use from API-key-backed access without recording credentials in the inventory.
  • Model selection: record whether the model is explicitly fixed, inherited from a workspace default or selected indirectly.
  • API route: for direct integrations, record whether the workload uses the Responses API, Chat Completions API or another documented supported route.
  • Version control: record whether the application uses a rolling model alias or a fixed snapshot, where documented and appropriate.
  • Owner and deadline: assign separate owners for surface migration and API evaluation when both apply.

This inventory is also a safeguard against accidental scope expansion. A consumer prompt workflow, a Codex coding task and a server-side Responses API application may all mention GPT-5.5, but they do not have the same controls, observability or change path. Evidence from one should not be presented as proof about the others.

The documented GPT-5.5 versus GPT-5.4 question

Once the surface issue is separated, the API decision becomes more precise: does GPT-5.5 produce enough measured value in a target workflow to justify its documented unit rates and configuration requirements relative to GPT-5.4? The model pages make GPT-5.4 the lower-priced documented baseline for this comparison, but price alone does not select the model.

OpenAI’s current model references list the same headline capacities for both models: a 1.05 million-token context window, a maximum output of 128,000 tokens, text and image input, and text output. They also list broadly matching endpoint and tool support. These shared specifications make controlled comparison feasible; they do not establish behavioural equivalence.

Equal context-window labels do not show that the models use long context with equal accuracy or consistency. Equal maximum-output labels do not show that both will produce the same length, structure or completion rate. Matching tool-support entries do not show equal tool selection, argument construction, stopping behaviour or recovery from tool errors. Those properties require evaluation in the intended workflow.

The model references also document a configuration difference that can invalidate a casual comparison: GPT-5.5 defaults to medium reasoning, whereas GPT-5.4 defaults to none. A test that invokes both models without recording reasoning.effort is not simply comparing model names; it is comparing different default reasoning configurations as well.

OpenAI’s GPT-5.5 guidance warns that higher reasoning effort can add latency and cost. It also warns that quality can regress where instructions conflict, stopping conditions are weak or tools are open-ended. This is vendor guidance about API configuration, not a universal prediction. It supports explicit tuning and measurement rather than automatically selecting the highest available effort.

A cost comparison should record the actual mix of uncached input, cached input and output for the same accepted cases. The later rate table provides the documented unit rates; this section defines the measurement record needed to apply them without turning a unit-price comparison into a claim about the completed-workflow bill.

Selected documented API comparison points
Attribute GPT-5.5 GPT-5.4 Decision significance
Standard input rate per million text tokens $5 $2.50 GPT-5.5 has twice the listed standard unit rate; completed-workflow cost still depends on actual use.
Standard cached-input rate per million text tokens $0.50 $0.25 Cache eligibility and hit rate must be measured rather than assumed.
Standard output rate per million text tokens $30 $15 Reasoning and output behaviour can change billed token totals, so the task bill is not determined by the rate alone.
Context window 1.05 million tokens 1.05 million tokens Equal listed capacity does not establish equal context utilisation or answer quality.
Maximum output 128,000 tokens 128,000 tokens A maximum is a capacity specification, not an expected output length.
Input and output modalities Text and image input; text output Text and image input; text output Representative image and mixed-input cases still require workflow testing.
Default reasoning effort Medium None Set and record reasoning explicitly for interpretable experiments.

Record long-context, regional-processing and service-mode conditions beside the observed token mix. These conditions belong in the same dated cost record as retries, tool calls and failed work, so an estimate can be reconciled with actual usage after a test.

Recommended cost rule: compare cost per accepted task, not price per token in isolation. For each test case, capture uncached input, cached input, output, retries, tool activity and whether the result passed the predefined quality gate. Include application and human-review costs only when the same accounting boundary is used for both models.

A simple token-cost worksheet can begin with the published rate categories, but it should be labelled an estimate. Do not apply the standard rates to long-context or regional-processing cases without checking the relevant conditions. Do not count a rejected answer as successful merely because it used fewer tokens. The denominator should be accepted work under a fixed evaluation policy.

How the evidence is weighted

Current model references establish documented API properties

The GPT-5.5 and GPT-5.4 model-reference pages are the primary sources for current API specifications in this review. OpenAI uses them to document model identifiers, modalities, context and output limits, reasoning settings, endpoints, tools, prices, knowledge cut-off information and snapshots. These pages are more appropriate for a September availability and configuration decision than historical launch wording.

A model-reference entry is evidence that OpenAI lists a capability; it is not proof that the capability performs adequately in a particular application. For example, listed Structured Outputs or tool support can establish interface availability, but not factual correctness, schema semantics, tool-choice reliability or safe production behaviour. Those require application-specific validation and human-led risk controls.

The Codex model notice establishes the surface retirement

The Codex models documentation is the controlling source in this review for the 14 October retirement statement and its API exclusion. Its wording should be preserved with both date and scope. Rephrasing it as a general GPT-5.5 shutdown would remove the most important qualifier and mislead API teams.

Because the notice mentions plan-dependent replacements, it can support a requirement to inspect each workspace. It cannot support a promise that a named replacement will be available to every user or behave as a drop-in substitute. Replacement validation should cover permissions, selected model, tools, latency, output quality and workflow-specific controls.

OpenAI’s selection guides support a method, not a winner

OpenAI’s “Using GPT-5.5” and model-selection guidance recommend representative examples, fresh baselines, same-input experiments and measured reasoning-effort tuning. OpenAI also advises retaining the lightest setting that meets the required quality threshold. This supports an evaluation process in which GPT-5.4 is tested as the lower-priced baseline and GPT-5.5 is retained only where measured results justify it.

“Representative” means that the evaluation set reflects the workload’s actual difficulty, formats, languages, context lengths, tool paths and failure consequences. A collection of convenient demonstration prompts is not representative merely because it is large. Include common cases, known failures, boundary cases and inputs where an incorrect action would require escalation or human review.

A fixed snapshot can improve reproducibility because the tested model version can be recorded rather than inferred from a moving alias. Snapshot availability and identifiers should be taken from the current model pages rather than guessed. Pinning a snapshot does not guarantee deterministic output or eliminate the need to retest application, tool and policy changes.

The April announcement is historical vendor evidence

OpenAI’s 23 April 2026 GPT-5.5 announcement provides launch context and a vendor-published evaluation table. It is not the source used here to determine current September availability or pricing. Launch-language claims can become outdated as product surfaces, model references and commercial conditions change.

The announcement’s evaluation table also mixes public, internal and differently qualified evaluations. Those results should retain OpenAI’s labels and methodological caveats if discussed. They are not an independent benchmark, do not establish universal superiority over GPT-5.4, and cannot substitute for a workload-specific comparison with disclosed prompts, versions, dates, scoring rules and outputs.

Evidence hierarchy for this review
Evidence type Appropriate use Inappropriate inference
Current Codex model documentation Determine the named surface retirement, effective date and API exclusion. Infer that GPT-5.5 is already retired on 30 September or retired from the API.
Current API model references Compare listed specifications, prices, settings, snapshots, endpoints and tools. Assume equal behaviour, access, limits or production suitability.
Current OpenAI model-selection guidance Design representative, same-input experiments and tune reasoning deliberately. Claim that either model wins every workload.
Historical OpenAI announcement Describe launch context and clearly labelled vendor evaluations. Treat launch claims as current availability evidence or independent proof.
Reader-run evaluation Measure target-workflow quality, latency, tokens, tools, retries and accepted-task cost. Generalise beyond the tested versions, inputs, settings and operating conditions.

Decision boundary for engineering teams

The decision framework below tests the lower-priced baseline and higher-priced alternative against the same acceptance contract.

Recommended evaluation record: preserve the model identifier or snapshot, date, endpoint, reasoning effort, system and developer instructions, tool definitions, stopping rules, input identifier, output, token counts, latency, retries, tool trace, grader decision and failure category. Without that record, a model-selection conclusion is difficult to reproduce or audit.

Safety and consequential decisions remain human-led. Matching endpoint support does not authorise autonomous legal, financial, health, security, employment or procurement actions. Where a model output can affect rights, access, money, safety or production systems, define human approval, source verification, rollback and incident handling independently of which model passes the quality test.

Before publication or deployment, recheck the live model pages and Codex notice. Confirm whether 14 October remains the effective date, whether the API exclusion remains explicit, whether model prices or long-context conditions have changed, and whether the required snapshot, endpoint and tools are available to the target account. Record deviations from the 30 September evidence baseline rather than silently updating one claim and leaving related assumptions unchanged.

GPT-5.5 and GPT-5.4 API feature matrix

The comparison below reflects OpenAI’s model references and selection guidance accessed on 30 September 2026. It does not establish that either model is universally more accurate, faster, safer, more reliable or more economical for a completed task.

The central comparison is narrower than the surrounding retirement news suggests. OpenAI’s Codex documentation schedules GPT-5.5 for retirement from ChatGPT, ChatGPT Work and Codex on 14 October 2026, while explicitly excluding the OpenAI API from that retirement. An API deployment using gpt-5.5 therefore requires an API model-selection review, not an assumption that the model has already become unavailable everywhere.

Token pricing and context limits visualised as layered analytical structures
Token pricing and context limits visualised as layered analytical structures.

Specification and configuration matrix

Documented GPT-5.5 and GPT-5.4 API properties, based on OpenAI documentation accessed 30 September 2026
Property GPT-5.5 GPT-5.4 Engineering significance
Primary model identifier (ID)A value used to distinguish one record, task, source or object from another. Open glossary entry gpt-5.5 gpt-5.4 Keep product-surface names out of API configuration. An API model ID is not interchangeable with a model selected in ChatGPT, ChatGPT Work or Codex.
API status at the evidence cut-off OpenAI’s retirement notice explicitly excludes the API. Documented in the current API model reference. This establishes documented API continuity, not guaranteed access for every account, region, endpoint or rate tier. Confirm access in the intended project before deployment.
Fixed snapshot OpenAI’s model reference lists snapshot information. OpenAI’s model reference lists snapshot information. This article does not reproduce the exact snapshot strings. Record the exact value from the live reference in the release manifest rather than guessing or deriving it from the alias.
Knowledge cut-off Specified on OpenAI’s model page. Specified on OpenAI’s model page. This article does not reproduce the exact cut-off dates. Copy them from the current references during publication or implementation review. A knowledge cut-off is not a guarantee that a fact is present, current or correct.
Context window 1.05 million tokens 1.05 million tokens Equal listed capacity does not demonstrate equal retrieval, instruction retention or context utilisation. Test documents at representative lengths and positions, including information near the beginning, middle and end.
Maximum output 128,000 tokens 128,000 tokens A maximum is a capacity boundary, not a recommended output target or a promise that every request can use the full amount. Input, reasoning and output configuration may interact with practical request limits.
Input modalities Text and images Text and images Matching modality labels do not prove equivalent optical interpretation, chart reading, extraction accuracy or image-related latency. Use the same authorised image set in comparative evaluation.
Output modalities Text Text Neither listing should be interpreted as documented native audio, image or video output support for this comparison.
Core API surfaces OpenAI lists broad support including the Responses API and Chat Completions API. OpenAI lists broadly matching endpoint support. Endpoint availability does not establish behavioural parity. Prefer a same-endpoint comparison so that transport and orchestration differences do not contaminate the model result.
Batch processing Listed as supported. The evidence reviewed for this article does not state Batch support; confirm it on the live GPT-5.4 model reference. Do not mix synchronous and Batch results in one latency or price comparison. They represent different delivery modes and may have different commercial conditions.
Structured outputs and tools OpenAI documents structured outputs and a broad tool-capable surface. OpenAI documents broadly matching endpoint and tool support. A listed tool is an available interface, not evidence of equal argument quality, tool selection, stopping behaviour or recovery after tool errors. Evaluate each tool loop used in production.
Default reasoning effort medium none An unmodified comparison tests different defaults. For an interpretable experiment, record the effective reasoning.effort value and compare both defaults and deliberately selected settings where supported.
Reasoning configuration OpenAI documents configurable reasoning effort. OpenAI documents reasoning options, with none as the default. Do not assume more reasoning is always better. OpenAI warns that higher GPT-5.5 effort can add latency and cost, and may regress quality when instructions conflict, stopping conditions are weak or tools are open-ended.
Standard input rate $5 per million text tokens $2.50 per million text tokens GPT-5.5’s listed standard input rate is twice GPT-5.4’s. This ratio applies to the published unit rates, not automatically to the final workload bill.
Standard cached-input rate $0.50 per million text tokens $0.25 per million text tokens The listed ratio is also two to one. Actual cached-input expenditure depends on whether requests qualify for caching and how much reusable prefix material is present.
Standard output rate $30 per million text tokens $15 per million text tokens The output-rate ratio is two to one. Output-heavy generation and reasoning behaviour can make output tokens materially important, so input-only estimates are incomplete.
Long-context threshold Higher long-context rates apply above 272,000 input tokens. Higher long-context rates apply above 272,000 input tokens. Do not price a request above the threshold using only the headline standard rates. The model pages state the threshold, but this article does not reproduce the higher rate schedule; retrieve it from the live model page before approving spend.
Regional processing OpenAI states a regional-processing uplift. OpenAI states a regional-processing uplift. This article does not state the uplift amount or every eligibility condition. Apply the current documented regional rate to the organisation’s actual processing configuration rather than treating standard pricing as universal.
Other price modes Batch, Flex, Fast mode and other commercial or processing conditions may alter applicable charges where available and selected. Keep these modes separate in the cost model. The standard rates above must not be presented as a complete price for every request path.

What the matching capacity figures do and do not show

Both model pages list a 1.05-million-token context window and a 128,000-token maximum output. Those figures establish matching documented ceilings, but they do not show how effectively either model uses information across the window. A model can accept a long request while still differing in citation accuracy, instruction retention, conflict resolution, latency, token consumption and its ability to recover relevant evidence from densely packed material.

Recommended evaluation procedure: divide long-context test cases into practical bands rather than testing only a short prompt and one request near the maximum. At minimum, include ordinary requests below the 272,000-token pricing threshold, requests close to that threshold and representative requests above it. Place decisive evidence at different positions and include plausible distractors. Score whether the answer uses the correct source passage, follows precedence rules and declines to invent missing evidence.

The 128,000-token output maximum should not become a default max_output_tokens setting without a workflow reason. Generous output bounds can permit unnecessary continuation, especially in tool-driven or weakly stopped tasks. Define an output contract, a stopping condition and a failure state. For extraction, classification and routing, a compact structured response may be more controllable than an open-ended narrative even though the model supports much longer output.

Equal text-and-image input labels also require task-level testing. A document workflow may contain scans, screenshots, charts, diagrams or text embedded in images. The relevant measure is not whether both references display “image input”, but whether the selected configuration extracts the required evidence reliably enough for the intended process. Consequential interpretations should remain subject to human verification against the source image.

Endpoint and tool support should be tested as behaviour, not inferred from a capability list

OpenAI’s current references describe broadly matching endpoint and tool support for GPT-5.5 and GPT-5.4. That reduces migration friction at the interface level, but it does not establish equivalent production behaviour. A tool-capable model must still choose the right tool, construct valid arguments, respect permissions, stop after satisfying the task and handle partial or erroneous tool results.

Recommended comparison rule: hold the orchestration layer constant. Send the same versioned system instructions, user inputs, tool definitions, output schema and stopping rules through the same API surface. If GPT-5.5 is tested through the Responses API while GPT-5.4 is tested through a materially different Chat Completions implementation, differences may come from orchestration rather than the model.

Tool evaluations should record more than final-answer acceptance. Capture tool calls per task, invalid arguments, unnecessary calls, repeated calls, permission failures, retries, incomplete responses and whether a human had to intervene. These events affect latency and total cost even when the final response eventually passes review.

Structured output support similarly constrains response shape rather than guaranteeing truth. A schema-valid object can contain an unsupported classification, an incorrect amount or a fabricated citation. Validation should therefore have two layers: machine validation for syntax and contract compliance, followed by domain checks or human review for semantic correctness where errors could affect users, finances, security, legal rights or operations.

Reasoning defaults create an asymmetric starting point

OpenAI documents medium as GPT-5.5’s default reasoning effort and none as GPT-5.4’s default. A simple alias swap therefore does not produce a configuration-neutral comparison. If a team sends identical requests without recording effective settings, it may compare a reasoning-enabled GPT-5.5 path against a no-reasoning GPT-5.4 path and then attribute every difference to the model family.

OpenAI’s GPT-5.5 guidance recommends establishing a fresh representative baseline and tuning reasoning effort through measurement rather than assuming a drop-in replacement. It also warns that higher effort can increase latency and cost, and can regress quality when prompts contain conflicting instructions, stopping criteria are weak or tools permit open-ended action. This is vendor guidance for the API, not a universal performance result or a statement about ChatGPT interface settings.

Recommended test design: begin with each model’s documented default because that represents an unmodified deployment. Then add controlled configurations relevant to the workload. Record the model alias or snapshot, reasoning.effort, prompt version, tool version, schema version and output limit for every run. Do not merge results from different settings into a single model-level score.

A higher-effort setting should be retained only if it improves a predeclared acceptance measure enough to justify its effect on latency, token use and failure behaviour. If an extraction task already meets its quality threshold at a lighter setting, more reasoning is not automatically beneficial. Conversely, if a complex planning task fails at a lighter setting, a measured higher-effort test may be warranted before rejecting the model.

Token rates: the documented ratio is exact, but the invoice ratio is not

For standard text-token rates, OpenAI lists GPT-5.5 at $5 per million input tokens, $0.50 per million cached-input tokens and $30 per million output tokens. GPT-5.4 is listed at $2.50, $0.25 and $15 respectively. Each GPT-5.5 unit rate is therefore exactly twice the corresponding GPT-5.4 rate.

That two-to-one relationship must not be rewritten as “GPT-5.5 costs twice as much per task”. Completed-task cost depends on the proportion of uncached input, cached input and output; the number of attempts; reasoning behaviour; tool calls; retries; incomplete responses; long-context treatment; regional processing; and any applicable Batch, Flex, Fast mode or contractual conditions. Models can also produce different token totals for the same task.

Recommended standard-rate estimate: calculate model-token cost separately for each recorded request using the applicable published rates:

estimated_token_cost =
  (uncached_input_tokens / 1,000,000 × input_rate)
+ (cached_input_tokens / 1,000,000 × cached_input_rate)
+ (output_tokens / 1,000,000 × output_rate)

This formula is a planning template, not a substitute for OpenAI billing records. It also excludes tool-specific charges, retries outside the measured request, engineering labour, review time and other operational costs. Requests above 272,000 input tokens require the current long-context schedule rather than an unqualified application of the headline rates.

For workflow comparison, divide total measured expenditure by accepted completed tasks, not by attempted requests. If a model needs fewer retries but uses more output tokens, or produces cheaper first attempts that require more human correction, per-request token cost can obscure the operational result. Keep model-token cost, tool cost and human-review effort as separate fields so that decision-makers can see why totals differ.

Recommended cost and operational record for each evaluation arm
Field Why it is required
Uncached input tokens Applies the normal input rate to the correct token category.
Cached input tokens Prevents cached and uncached input from being priced as if they were identical.
Output tokens Captures the higher-priced output component and differences in response length.
Input length band Identifies requests that cross the 272,000-token long-context threshold.
Processing region and price mode Prevents standard, regional, Batch, Flex or Fast mode conditions from being mixed.
Tool calls and tool charges Separates model-token economics from external or platform tool activity.
Retries and incomplete responses Measures cost incurred before an accepted completion.
Human review or correction time Supports a total-cost-of-ownership comparison without pretending labour is a token charge.
Accepted completed tasks Provides the denominator for cost per accepted result.

Snapshot and knowledge cut-off controls

OpenAI’s references list snapshot information and knowledge cut-offs for both models. Those fields exist on the model pages, but this article does not reproduce their exact values. Before publication or implementation, copy the precise snapshot identifiers and cut-off dates from the live pages into the deployment record, along with the date on which they were verified.

A fixed snapshot improves experimental traceability because it reduces ambiguity about which documented model revision produced an output. It does not by itself guarantee full reproducibility: API infrastructure, tools, external data, stochastic behaviour and surrounding application code can still change. Preserve representative inputs, configuration, tool responses where permitted, expected outputs, grading rules and run dates alongside the snapshot.

A knowledge cut-off should be treated as a boundary on the model’s training knowledge, not as a certification of factual coverage. For current or consequential information, provide an authorised source, require evidence-linked answers and verify the result. Neither model should be allowed to make irreversible decisions merely because its listed cut-off is newer or appears more suitable.

Selection rule for this matrix

GPT-5.4 is the lower-priced documented baseline: its three listed standard token rates are half the corresponding GPT-5.5 rates, while the references list the same context capacity, maximum output and input/output modalities, plus broadly matching endpoints and tools. This makes it a rational baseline for testing, not a predetermined choice.

GPT-5.5 should be retained or introduced only where a representative evaluation shows that its accepted quality, tool behaviour or other workflow result justifies its higher unit rates and configuration requirements. OpenAI’s model-selection guidance supports same-input experiments and retaining the simplest passing configuration that meets the quality threshold. The quality threshold, risk tolerances and review requirements remain the deploying organisation’s decisions.

Recommended decision rule: choose the least costly model-and-reasoning configuration that passes the predeclared quality, safety, latency and operational acceptance gates on representative inputs. Escalate to GPT-5.5 only where measured gains survive review after token, tool, retry and human-correction costs are included.

This rule does not make the API lifecycle question disappear. Teams should separately inventory ChatGPT, ChatGPT Work and Codex configurations affected by the scheduled 14 October 2026 retirement. API continuity does not update those product surfaces, while the product-surface retirement does not prove that direct API calls to gpt-5.5 are retired. Treat them as related but distinct workstreams.

Representative evaluation and migration decisions shown as a measured comparison path
Representative evaluation and migration decisions shown as a measured comparison path.

Build a representative evaluation before choosing a model

Recommended framework: define an acceptance contract, freeze the configurations, run both models on the same representative cases, record complete workflow telemetry, adjudicate quality without model-name bias where practical, and calculate accepted-work cost rather than token price alone. This is a proposed evaluation method, not a claim that either model will win a particular workload.

1. Define the production decision and acceptance contract

Start with one bounded decision. Examples include choosing a model for contract-clause extraction, support-ticket classification, repository question answering, or a tool-using research step. Avoid combining unrelated workloads into one average because a strong result on routine extraction can conceal unacceptable failures on escalation-sensitive cases.

Write the quality threshold before inspecting comparative outputs. The contract should identify mandatory output fields, factual-support rules, permitted sources, tool constraints, escalation conditions, and disqualifying errors. For a document-extraction workflow, a record might pass only when every extracted value is supported by the supplied document, required fields conform to the schema, uncertainty is represented as instructed, and no value is inferred from outside knowledge.

Separate hard gates from optimisation measures. A hard gate is a condition that must be met, such as no unsupported account action, valid structured output, or correct handling of a designated escalation category. Optimisation measures—latency, token use, or stylistic preference—matter only after the candidate passes the hard gates. This prevents a cheaper aggregate result from compensating for an unacceptable consequential error.

Evaluation layer Example measure Recommended decision treatment
Safety or policy gate Prohibited action attempted; required escalation omitted Fail the case or configuration; require human investigation
Contract validity Schema validity, required fields, citation presence Fail or retry according to the predeclared production policy
Task quality Correct classification, supported answer, complete extraction Compare against a predeclared acceptance threshold
Operational performance End-to-end latency, tool calls, retries, incomplete responses Compare distributions and failure modes, not only averages
Economics Cost per attempted case and cost per accepted case Evaluate after quality and safety gates have been applied

2. Construct a case set that reflects production risk

A representative evaluation is not a random collection of convenient prompts. Build the set from authorised production-like examples, incident records, support escalations, known edge cases, and deliberately constructed boundary cases. Remove or minimise personal, confidential, regulated, and security-sensitive information according to organisational policy. Synthetic cases can broaden coverage, but they should not be presented as proof of production prevalence.

Stratify cases by factors likely to change model behaviour. Useful strata include short versus long input, clean versus ambiguous instructions, single versus multiple documents, relevant evidence near the beginning versus near the end, no-tool versus multi-tool paths, expected abstention, conflicting evidence, malformed tool results, and tasks requiring several dependent steps. Record the stratum for every case so aggregate results can be decomposed.

Include ordinary high-volume cases as well as rare high-consequence cases. Sampling only difficult examples can misrepresent routine operating cost, while sampling only common examples can miss the failures that determine whether deployment is acceptable. Weighting should reflect the explicit decision: production-frequency weighting estimates expected operations, whereas risk-weighted reporting gives greater visibility to consequential cases. Report both separately rather than silently combining them.

Reserve a holdout set that prompt authors do not repeatedly inspect during tuning. Use a development set to refine instructions and a frozen holdout for the final comparison. Repeatedly changing prompts after examining holdout failures turns the holdout into another development set and makes the resulting acceptance estimate optimistic.

3. Freeze prompts, snapshots and runtime controls

Record the exact model identifier or documented snapshot where appropriate, evaluation date, endpoint, system and developer instructions, user input, tool definitions, response-format contract, reasoning effort, maximum-output setting, timeout policy, retry policy, and any truncation or preprocessing. Fixed snapshots improve reproducibility, although they do not remove variation caused by infrastructure, tools, external data, or non-deterministic generation.

Run paired cases with identical task inputs and equivalent surrounding conditions. If a prompt must be changed for one model, classify that as a separately tuned configuration rather than a direct same-prompt comparison. Both views can be useful: the same-prompt test measures drop-in behaviour, while the tuned test measures the best configuration the team is prepared to maintain.

Do not omit reasoning.effort unintentionally. At minimum, test the actual configuration proposed for each model and document the asymmetry. If the purpose is to understand reasoning sensitivity, add explicitly named variants rather than treating every effort level as the same candidate. OpenAI recommends fresh representative baselines and measured tuning instead of assuming GPT-5.5 is a drop-in replacement.

Keep production retry and fallback logic visible. A harness that retries indefinitely until it obtains a valid answer measures a different service from an application that permits one retry and then escalates. The evaluation should enforce the proposed timeout, retry ceiling, tool-call ceiling, and human-review route.

Measure six dimensions at the workflow level

Accepted quality, not surface fluency

Quality scoring should be task-specific and evidence-based. Prefer deterministic checks where they are genuinely meaningful: schema validation, exact identifiers, arithmetic reconciliation, required citations, and tool-argument validation. Use qualified human reviewers for ambiguity, factual support, instruction compliance, and consequential judgements. Automated graders can assist triage, but their outputs should not be treated as ground truth without validation against human decisions.

Use a documented rubric with examples of pass, partial pass, and fail. Reviewers should record the error category as well as the score: unsupported assertion, omitted requirement, wrong source, invalid tool argument, premature stopping, over-escalation, or formatting failure. Error categories reveal whether a configuration can be repaired through prompt or tool design; a single aggregate score cannot.

Blind model labels during review where operationally feasible. Randomise output order and preserve a case identifier so reviewers do not know whether an answer came from GPT-5.5 or GPT-5.4. Blinding does not remove rubric ambiguity, but it reduces expectation bias. Resolve material reviewer disagreements through a recorded human adjudication process.

Report quality by stratum and with raw counts. A configuration that passes 95 of 100 cases may still be unsuitable if all five failures occur in the same escalation-sensitive category. Do not declare a winner from a small numerical difference without considering sample size, disagreement, repeated-run variation, and the practical importance of the affected cases.

End-to-end latency rather than model-call timing alone

Measure latency from the application’s receipt of a valid request to delivery of an accepted result or escalation. Record model time, tool time, queueing, retries, parsing, validation, and any fallback separately. User-perceived or job-completion latency can rise even when the initial model response is quick if the configuration makes more tool calls or produces more invalid outputs.

Report a distribution rather than only a mean. Recommended summary fields include median, 90th percentile, 95th percentile, maximum observed within the run, timeout count, and time to accepted completion. These are reporting choices, not promised service levels. Run comparable loads and avoid inferring account-wide rate behaviour from a lightly loaded test.

Warm caches, external-service variability, regional routing, and concurrent demand can affect results. Record concurrency, request schedule, cache status where observable, and tool-service conditions. Alternate or randomise model order so one candidate is not always tested during a systematically quieter period.

Token use and context utilisation

Capture input, cached-input, and output tokens from the response metadata available to the application. Keep reasoning-related consumption visible where the API reports it, and avoid estimating it from visible answer length. A concise answer is not necessarily a low-token execution, particularly when reasoning settings differ.

For long-context cases, record total submitted tokens, placement of decisive evidence, amount of irrelevant material, and whether the accepted answer used the correct evidence. The 1.05 million-token listed window for both models does not show how either behaves near the limit. OpenAI’s model pages also state higher long-context rates above 272,000 input tokens, so those cases require the applicable documented price tier rather than the standard input rate.

Separate cache-eligible input from cache hits actually billed at the cached-input rate. A theoretical reusable prefix is not a realised saving. Tests should reproduce the application’s expected request structure and repetition pattern rather than assuming every shared instruction or document will receive cached pricing.

Retries, failures and recovery

Define a retry taxonomy before the run. Transport failures, timeouts, rate-limit responses, incomplete model responses, invalid structured output, failed tool calls, and semantically unacceptable answers are different events. Retrying every failure with the same policy can hide quality defects and inflate both latency and cost.

For each attempted case, record the number of model requests, tool invocations, validation failures, retry reason, final disposition, and whether a human had to intervene. Cap retries as production would. If a second attempt changes the prompt, reasoning effort, model, or tool availability, log it as a fallback step rather than another identical sample.

Evaluate recovery correctness as well as recovery rate. A model that notices a tool error but invents a result has not recovered. A safe recovery may be to retry a permitted transient operation, request missing information, return a bounded failure, or route the case to a person. Which response is correct depends on the application contract.

Tool choice, arguments and stopping behaviour

Tool evaluation should inspect the entire trajectory. Score whether the model selected an allowed tool, supplied valid and semantically correct arguments, respected authorisation boundaries, interpreted the result correctly, avoided redundant calls, and stopped when the task was complete. A valid function-call schema does not prove that the requested operation was appropriate.

Use deterministic mock tools for the first comparison where possible. Mocks can return controlled success, empty, malformed, delayed, conflicting, and permission-denied results. This isolates model behaviour from changing external systems and makes failures reproducible. Follow with a bounded integration test because mocks cannot establish real-service reliability.

Set explicit call and loop limits. Open-ended tools combined with weak stopping conditions are among the situations OpenAI identifies as capable of producing worse outcomes at higher GPT-5.5 reasoning effort. The appropriate control is not to assume lower or higher effort is always preferable, but to test the proposed effort with finite tool budgets and clearly defined completion conditions.

Total cost per accepted outcome

Calculate cost from observed input, cached-input and output volumes for accepted cases, then add retries and measurable tool charges. The documented rate table appears earlier; this stage applies it to the evaluation record rather than repeating it.

Calculation: calculate each case from the rates and conditions applicable when the evaluation is run, then add model retries and measurable tool charges. Do not apply the standard rate to inputs above the documented 272,000-token long-context threshold. Account separately for any documented regional-processing uplift and for Batch, Flex, Fast mode, or other processing conditions actually used; do not assume their terms.

case_token_cost =
  input_tokens × applicable_input_rate
  + cached_input_tokens × applicable_cached_input_rate
  + output_tokens × applicable_output_rate

case_workflow_cost =
  sum(model-request token costs)
  + metered tool costs
  + other directly attributable processing costs

cost_per_accepted_case =
  total workflow cost for the evaluated cohort
  ÷ number of cases meeting the acceptance contract

Use consistent token units when applying per-million-token rates. Keep labour and infrastructure separate unless the organisation has a defensible allocation method. Human review time, engineering maintenance, observability, incident handling, and vendor-tool charges may matter to total cost of ownership, but invented hourly values would make the comparison less reliable. Record actual internal assumptions and run sensitivity ranges instead.

Cost per accepted case is usually more decision-relevant than cost per first request. If one configuration uses fewer first-pass tokens but requires more retries or human corrections, its workflow cost can exceed what the initial call suggests. Conversely, higher token rates can sometimes be offset by fewer attempts or shorter outputs, but that possibility must be measured rather than presumed.

Use a configuration ladder instead of a single-pass comparison

Experiment design: begin with GPT-5.4 at its documented default reasoning setting as the lower-priced baseline, then test explicitly configured alternatives only where the baseline misses the acceptance contract. Include GPT-5.5 at the actual proposed reasoning effort, remembering that its documented default is medium. This sequence follows OpenAI’s general guidance to retain the lightest setting that meets the quality threshold; it does not prejudge which candidate will pass.

  1. Baseline: run GPT-5.4 with the frozen production candidate prompt and explicit runtime controls.
  2. Diagnose: classify failures by task stratum, error type, tool stage, and retry cause.
  3. Repair the workflow first: remove conflicting instructions, strengthen stopping conditions, constrain tools, or improve source formatting when the defect is model-independent.
  4. Run the GPT-5.5 candidate: use the same cases and declare its reasoning effort rather than relying on an undocumented assumption.
  5. Tune narrowly: test only changes supported by a failure hypothesis, and keep tuned configurations distinct from same-prompt results.
  6. Confirm on the holdout: freeze the selected candidates and run the untouched cases under representative operational conditions.
  7. Approve conditionally: document where the selected model is permitted, its fallback, its review requirements, and the metrics that trigger reassessment.

A mixed deployment can be more defensible than a single global choice. For example, GPT-5.4 might handle cases that pass deterministic prechecks, while a narrowly defined difficult stratum is evaluated for GPT-5.5. That is a recommendation to test routing, not a claim that either model is suited to a named task. Routing adds classification errors, maintenance, and observability requirements, which must be included in the evaluation.

Interpret OpenAI’s launch results as vendor-published evidence

OpenAI’s 23 April 2026 announcement contains a vendor-published launch evaluation table. It is historical product-announcement evidence, not an independent benchmark and not current availability evidence. The table mixes public and internal evaluations and uses differently qualified settings, so its results should retain OpenAI’s labels, conditions, and methodological caveats rather than being compressed into a universal superiority claim.

Launch results can help identify test categories worth reproducing, but they cannot substitute for a workload-specific evaluation. Production prompts, tools, document distributions, stopping rules, reasoning settings, safety requirements, and acceptance thresholds may differ from OpenAI’s evaluated conditions. Without disclosed matching prompts, versions, dates, methodology, outputs, and spend assumptions, numerical comparison with an internal run would not be like-for-like.

No selection verdict should be derived from the launch table alone. A defensible review can state that OpenAI reported particular results under its evaluated settings, while withholding a production judgement until the organisation has run representative paired tests. Current API model references and current Codex documentation should govern availability and configuration claims, not historical launch wording.

Apply a documented selection gate

Select a configuration only if it passes every mandatory safety and contract gate, meets the predeclared quality threshold on the frozen holdout, remains within the accepted latency and failure envelope, and has a supportable cost per accepted outcome. If both candidates pass, prefer the lighter and lower-cost configuration unless a material, measured benefit justifies the alternative. If neither passes, revise the workflow, narrow the task, add human review, or decline deployment rather than lowering a consequential quality threshold after seeing the results.

Observed result Recommended decision
GPT-5.4 passes all gates and GPT-5.5 adds no material workflow benefit Retain GPT-5.4 as the documented lower-priced baseline.
Both pass, but GPT-5.5 materially improves a predeclared measure Calculate whether the measured improvement justifies higher applicable cost and tuning overhead.
GPT-5.4 fails a bounded stratum and GPT-5.5 passes it Consider a scoped GPT-5.5 deployment or evaluated routing rule; do not generalise the result to all traffic.
GPT-5.5 improves quality but breaches latency or cost limits Test a lower explicit reasoning effort, tighter tools, or narrower use; reject if the full contract remains unmet.
Results vary materially across repetitions or reviewers Increase investigation, refine the rubric, and avoid a winner claim until uncertainty is operationally acceptable.
Neither configuration passes a mandatory gate Do not deploy the model-only workflow; redesign it or preserve human handling.

Archive the evaluation manifest, case-set version, model snapshots or identifiers, prompts, tool schemas, raw outputs, telemetry, reviewer decisions, pricing assumptions, and approval record. Re-run the relevant holdout when changing a model snapshot, reasoning effort, prompt, endpoint, tool definition, preprocessing step, retry policy, or material data source. Live documentation should also be rechecked because listed availability, features, prices, and conditions can change.

This API decision remains separate from the announced product-surface retirement. As documented on 30 September 2026, the 14 October retirement concerns GPT-5.5 in ChatGPT, ChatGPT Work, and Codex and explicitly excludes the API. An organisation may therefore select GPT-5.5 for a justified direct-API workflow while separately updating affected ChatGPT- or Codex-based configurations before the prospective retirement date. API continuity does not guarantee identical access across accounts, regions, endpoints, or rate tiers, so deployment owners must verify their own current access.

Conditional recommendations for model selection and retirement readiness

The practical recommendation is to separate two decisions that happen to involve the same model name. First, owners of ChatGPT, ChatGPT Work and Codex workflows should prepare for GPT-5.5’s scheduled 14 October 2026 product-surface retirement. Second, owners of direct OpenAI API workloads should decide whether GPT-5.5 or GPT-5.4 provides the better production configuration for each workload. OpenAI’s documentation accessed on 30 September 2026 says the retirement notice does not retire GPT-5.5 from the API, so a surface migration should not be used as evidence that an API model change is required.

A conservative default for new API evaluations is to use GPT-5.4 as the lower-priced documented baseline, then admit GPT-5.5 only when a representative evaluation shows enough additional value to justify its higher unit rates and configuration requirements. This is a recommendation, not a claim that GPT-5.4 is universally preferable. GPT-5.5 may be the better choice for a particular workload, but that conclusion should be based on accepted outputs, latency, token consumption, tool behaviour, retries and total workflow cost rather than the model name or a launch table.

Scenario 1: a new API workload without an established baseline

Recommendation: use the lower-priced documented configuration as the starting point, then keep only a configuration that passes the defined quality gate.

Build the first evaluation set from real or realistically transformed production cases, including ordinary requests, long inputs, ambiguous instructions, malformed tool results, expected refusals and cases requiring human escalation. Record which cases are acceptable under a written rubric. If GPT-5.4 passes, retain it unless another measured requirement supports testing GPT-5.5. If it fails, test targeted prompt or workflow corrections before increasing reasoning effort or changing models; otherwise the experiment cannot distinguish model capability from an avoidable interface defect.

When GPT-5.5 is introduced, do not copy the GPT-5.4 configuration and assume the comparison is controlled. OpenAI documents different default reasoning settings: GPT-5.5 defaults to medium reasoning, while GPT-5.4 defaults to none. Set and record reasoning.effort explicitly where the API and selected configuration support it. This prevents a nominal model comparison from silently becoming a comparison between different reasoning budgets.

Scenario 2: an existing GPT-5.4 workload misses a material quality threshold

Recommendation: test GPT-5.5 against the failed strata rather than rerunning only an aggregate sample. For example, if GPT-5.4 passes routine extraction but fails multi-document conflict resolution, construct a labelled stratum for that failure mode and retain an unchanged general set to detect regressions. A model should not be promoted merely because it improves the difficult subset if it introduces unacceptable errors, latency or tool loops elsewhere.

Use an acceptance rule defined before inspecting the results. A suitable template is: promote GPT-5.5 only if it meets every safety and correctness floor, improves the designated failure stratum by the organisation’s pre-approved margin, remains within the latency budget, and keeps cost per accepted outcome below the approved ceiling. The organisation must choose those thresholds according to operational risk; OpenAI’s documentation does not supply a universal pass mark.

Test more than one GPT-5.5 reasoning setting where justified, but stop once the simplest configuration meets the acceptance contract. OpenAI warns that increasing GPT-5.5 reasoning effort can add latency and cost, and that conflicting instructions, weak stopping conditions or open-ended tools can cause quality regressions. Higher effort should therefore be treated as an experimental variable rather than a monotonic quality control.

Scenario 3: an established GPT-5.5 API workload already passes

Recommendation: do not migrate solely because GPT-5.5 is scheduled to leave ChatGPT, ChatGPT Work and Codex. The current retirement notice explicitly excludes the API. Instead, run GPT-5.4 as a challenger because its listed standard token rates are half those of GPT-5.5 across input, cached input and output. A challenger evaluation can identify workloads for which GPT-5.4 preserves accepted quality at a lower unit rate.

Migration should remain conditional on total cost and operational behaviour. A GPT-5.4 call that consumes more tokens, triggers more retries, produces more rejected outputs or requires more human correction may not be cheaper per accepted result. Conversely, if both configurations pass and their operational profiles remain within tolerance, the lower-priced configuration is the defensible default.

Keep the incumbent GPT-5.5 configuration available for rollback while the GPT-5.4 challenger is evaluated in a non-user-visible comparison or introduced to a limited share of eligible production traffic, subject to the organisation’s data-handling and deployment controls. Compare the same eligible inputs, but avoid duplicate consequential actions: tool calls that send messages, modify records, execute purchases or affect user access should be simulated, stubbed or restricted to a non-production environment during comparison.

Scenario 4: long-context or tool-heavy processing

Recommendation: treat the matching listed context window, maximum output and broad tool support as eligibility checks, not selection evidence. Both model references list a 1.05-million-token context window and a 128,000-token maximum output, with text and image input and text output. Those figures do not demonstrate that the models retrieve, prioritise or reason over long inputs equally well.

Create tests that place decisive evidence at different positions, include distractors and contradictions, and verify whether the output cites or uses the correct material. For tool-heavy workflows, log tool selection, argument validity, repeated calls, stopping behaviour, recovery after tool errors and any need for human intervention. OpenAI’s warning about open-ended tools and weak stopping conditions makes unbounded agent loops a specific evaluation risk for higher-reasoning GPT-5.5 configurations.

Both model pages state that higher long-context rates apply above 272,000 input tokens. Do not apply the standard input rate to an entire forecast without checking the current long-context charging rules. Chunking, retrieval, summarisation and caching may change the bill, but they may also omit evidence or alter behaviour; each optimisation requires an accuracy check against full-context cases.

Scenario 5: regulated, security-sensitive or consequential workflows

Recommendation: keep the decision human-led and require a risk-specific review independent of the model comparison. Neither a documented feature nor a successful aggregate evaluation proves legal compliance, security suitability, fairness or fitness for consequential decisions. Sensitive actions should retain authorised human approval, auditable evidence and a tested fail-safe path.

Use redacted or authorised evaluation data, minimise retained content, and review account, region and processing requirements before deployment. The retirement notice’s statement that the API is unaffected does not guarantee identical regional availability, account access, rate tier or endpoint access. Procurement and security teams should verify the actual configuration rather than relying on a generic model-page capability.

Migration implications differ by product surface

Surface Documented implication Required owner action What not to infer
ChatGPT GPT-5.5 is scheduled to retire from this surface on 14 October 2026. Inventory saved workflows, instructions and human operating procedures that explicitly depend on GPT-5.5; test the replacement actually available to the relevant account. Do not assume the model was already unavailable on 30 September 2026 or that a particular replacement is guaranteed for every plan.
ChatGPT Work The same scheduled retirement applies, while replacement access can depend on plan and workspace conditions. Have workspace administrators confirm permissions, available replacements and affected shared procedures; obtain owner sign-off before changing governed workflows. Do not equate workspace availability with consumer ChatGPT or direct API access.
ChatGPT-authenticated Codex The Codex documentation includes this product surface in the scheduled retirement. Review model selections, coding instructions, tool permissions and approval gates; re-run repository-specific tests before adopting an available replacement. Do not assume API continuity preserves a Codex user-interface selection.
API-key Codex integration Its behaviour depends on the integration’s model selection and the distinction between Codex product access and API access. Inspect configuration, credentials, model identifiers and execution path rather than classifying it by the word “Codex” alone. Do not assume every Codex-labelled workflow is retired or unaffected without tracing how it authenticates and invokes the model.
Direct OpenAI API OpenAI’s notice says the 14 October retirement does not apply to the API. Continue normal lifecycle monitoring and evaluate GPT-5.5 against GPT-5.4 according to workload evidence. Do not interpret “unaffected” as guaranteed access across all accounts, regions, endpoints or rate tiers.

Maintain separate inventory fields for product surface, authentication method, model identifier, snapshot, endpoint, reasoning setting and business owner. A single row labelled “uses GPT-5.5” is insufficient because it cannot show whether the scheduled retirement applies or whether the workload is a direct API deployment that merely needs an economic review.

For product-surface retirement work, preserve representative prompts and expected outcomes but avoid assuming that settings transfer exactly between interfaces. For API migration, pin and record a documented snapshot where appropriate, then test the replacement snapshot or alias before changing production traffic. Fixed snapshots improve reproducibility, although they do not remove the need to monitor surrounding tools, prompts, data and service conditions.

A cost worksheet for comparable production estimates

The following formulas are a recommended worksheet, not a quotation or billing guarantee. Populate them with measured token counts and the current rates applicable to the account, processing mode, region, context length and tools. The standard rates documented on 30 September 2026 can seed ordinary cases, but long-context, regional processing and other charging conditions must be represented separately.

Standard token cost per completed attempt

token_cost =
  (uncached_input_tokens / 1,000,000 × input_rate)
+ (cached_input_tokens / 1,000,000 × cached_input_rate)
+ (output_tokens / 1,000,000 × output_rate)

For standard-rate rows, insert GPT-5.5 rates of $5 input, $0.50 cached input and $30 output per million tokens, or GPT-5.4 rates of $2.50, $0.25 and $15 respectively. Keep cached and uncached input separate. Applying the cached rate to an assumed cache hit rather than an observed eligible hit will understate expected cost.

Expected cost per submitted task

expected_task_cost =
  initial_attempt_cost
+ (retry_probability × average_retry_cost)
+ expected_tool_charges
+ expected_other_processing_charges

This calculation should include model retries caused by timeouts, invalid structured results, tool failures, policy handling or application-level rejection. “Other processing charges” is a worksheet category for applicable documented conditions such as regional processing, long context or selected service modes; it is not permission to invent a blanket multiplier.

Cost per accepted result

cost_per_accepted_result =
  total_model_and_tool_cost_for_evaluation_set
  / number_of_results_that_pass_the_acceptance_rubric

This is more decision-useful than cost per call when outputs can fail review. If 1,000 calls are inexpensive but many require reruns or human correction, the apparent token saving may not survive at the accepted-result level. Report the number of rejected, retried, escalated and abandoned cases beside this figure so that a low denominator cannot be overlooked.

Human-review and operational cost

human_review_cost =
  review_hours × approved_loaded_hourly_cost

total_cost_of_ownership =
  token_cost
+ tool_and_processing_cost
+ human_review_cost
+ infrastructure_cost
+ monitoring_and_incident_cost

The loaded hourly cost and operational categories are organisation-supplied assumptions, not OpenAI prices. Use the same accounting boundary for both models. Counting reviewer time for one configuration while excluding it from the other invalidates the comparison.

Break-even value for GPT-5.5

incremental_cost =
  GPT-5.5_total_cost - GPT-5.4_total_cost

required_incremental_value =
  incremental_cost + required_risk_margin

Approve GPT-5.5 only if the measured value of additional accepted outcomes, avoided remediation or another authorised business metric meets the required incremental value without breaching safety, latency or reliability floors. Where value cannot be credibly monetised, apply a rule-based decision: both models must pass mandatory controls, then select the least costly passing configuration.

Migration controls and rollback evidence

  1. Freeze the incumbent record. Store the model identifier or snapshot, endpoint, prompt version, reasoning setting, tool definitions, timeout policy and evaluation date. Remove credentials and personal data from evaluation artefacts unless their inclusion is authorised and necessary.
  2. Run the challenger on the same eligible cases. Preserve input parity while preventing duplicate external actions. Record failures rather than silently resubmitting until a usable answer appears.
  3. Compare by risk stratum. Separate routine, difficult, safety-sensitive, long-context and tool-dependent cases. Aggregate averages can conceal a severe regression in a small but consequential category.
  4. Run a limited production test under explicit limits. If offline evaluation passes, route a controlled share of suitable production traffic to the challenger. Define stop conditions for quality, latency, cost, tool errors and incident signals before starting.
  5. Retain a tested rollback. Rollback should restore the complete known configuration, not only the model name. A model reversal paired with changed prompts or tool schemas is not a controlled restoration.
  6. Obtain accountable approval. Engineering, product, security and other relevant owners should approve according to the workflow’s risk. The model’s output should not approve its own deployment.

Material risks to record in the decision log

  • Surface confusion: teams may prematurely remove a functioning API integration or fail to update a retiring ChatGPT or Codex workflow because both use the GPT-5.5 name.
  • Default-setting bias: GPT-5.5’s medium default reasoning and GPT-5.4’s none default can produce an uncontrolled comparison if reasoning is not explicitly recorded.
  • Invoice extrapolation: exactly doubled standard token rates do not prove a doubled bill because token mix, cache use, long context, tools, retries, processing conditions and acceptance rates differ.
  • Capability-list bias: matching listed context, output, endpoint and tool support can be mistaken for equal quality, latency or reliability.
  • Launch-evidence overreach: OpenAI’s April 2026 announcement contains vendor-published evaluations with differing qualifications. It is historical product evidence, not an independent production benchmark or current availability record.
  • Open-ended execution: weak stopping conditions or permissive tools can increase calls, latency and risk, particularly when reasoning effort is raised without workflow controls.
  • Alias drift: an unpinned model alias may complicate reproduction. Record snapshots where appropriate and re-evaluate when a selected model or surrounding dependency changes.
  • False economy: selecting solely on unit rates can move cost into retries, review, incidents or failed outcomes.

Frequently asked questions

Must every GPT-5.5 API application migrate before 14 October 2026?

No. According to OpenAI’s documentation accessed on 30 September 2026, the scheduled retirement applies to ChatGPT, ChatGPT Work and Codex and explicitly excludes the API. API owners should still verify live availability and their own account conditions, but the cited notice does not establish an API retirement deadline.

Is GPT-5.4 the recommended replacement for every retiring surface?

No. The available replacement can depend on the product surface, plan, workspace and rollout state. This review uses GPT-5.4 as a lower-priced API baseline; it does not guarantee that GPT-5.4 is the replacement presented in ChatGPT, ChatGPT Work or Codex.

Does GPT-5.5 cost exactly twice as much as GPT-5.4?

Its listed standard rates are exactly twice GPT-5.4’s rates for input, cached input and output tokens. That ratio must not be translated directly into a claim about the final bill. Token volumes, cache hits, output length, long-context pricing, regional processing, tools, retries and rejected outputs can change total cost.

Can the higher-priced model be assumed to produce better answers?

No. Price is not a quality guarantee. OpenAI’s launch evaluations are vendor-published results under specified or differently qualified conditions, not proof that GPT-5.5 wins every production workload. Run a representative evaluation using the target prompt, tools, data shape and acceptance rubric.

Are matching context windows evidence that migration is low risk?

No. Matching listed capacity means both models are documented to accept the same maximum context scale, subject to current service conditions. It does not establish equal attention to distant evidence, resistance to distractors, latency, token use or output quality.

Should reasoning effort always be increased when quality is insufficient?

No. First check instruction conflicts, missing context, tool schemas, stopping rules and the acceptance contract. OpenAI warns that higher GPT-5.5 reasoning effort can increase latency and cost and may regress quality in poorly bounded workflows. Increase it only as a measured experimental step.

What is the minimum evidence for a production decision?

A defensible minimum is a dated representative case set, written acceptance rubric, recorded model and snapshot identifiers, explicit reasoning settings, prompt and tool versions, accepted-output counts, end-to-end latency, token use, retries, tool failures, cost assumptions and human approval. Higher-risk deployments require additional security, privacy, legal or domain review.

Final decision rules

  • For a new API workload, use GPT-5.4 as the documented lower-priced baseline and test GPT-5.5 only where the baseline misses a material requirement.
  • For a passing GPT-5.5 API workload, test GPT-5.4 as a challenger; migrate only if it preserves required quality and controls while improving the chosen cost or operational objective.
  • For a failing GPT-5.4 workload, admit GPT-5.5 only after a same-input evaluation shows a meaningful improvement and the complete configuration remains inside latency, cost and risk limits.
  • For ChatGPT, ChatGPT Work and Codex, complete a separate retirement inventory before 14 October 2026 and test the replacement actually available to each account or workspace.
  • For every surface, recheck current documentation and access before implementation because the 30 September 2026 evidence cut-off cannot guarantee later availability, plan conditions or service configuration.

Keep the API evaluation separate from the scheduled product-surface retirement programme.

Scope and evidence boundary: The documented 14 October 2026 retirement, prospective as of 30 September 2026, affects GPT-5.5 in ChatGPT, ChatGPT Work and Codex; it does not apply to the OpenAI API, so this is not an API retirement. OpenAI documents GPT-5.5 standard rates of $5 input, $0.50 cached input and $30 output per million tokens, and GPT-5.4 rates of $2.50 input, $0.25 cached input and $15 output per million tokens. Both model pages document a 1,050,000-token context window and up to 128,000 output tokens. This comparison is documented decision support, not an independent benchmark; use representative evaluation tasks before changing a workload.

Use a two-track inventory before making any model decision

The 14 October notice and the API comparison require separate workstreams. A product-surface inventory determines which ChatGPT, ChatGPT Work, or Codex configurations need attention before the scheduled retirement. An API inventory determines whether applications that select gpt-5.5 should remain unchanged, be tested against GPT-5.4, or move only after representative evaluation. Combining these workstreams can create a false deadline for direct API deployments or, in the other direction, leave a retiring product-surface configuration unaddressed.

Product-surface inventory checklist

  • Record whether each workflow runs in ChatGPT, ChatGPT Work, ChatGPT-authenticated Codex, API-key Codex, or the direct OpenAI API.
  • Identify the account, plan, workspace and permissions governing access without assuming that one surface’s availability applies to another.
  • Locate saved configurations, instructions, automations and documentation that explicitly select GPT-5.5.
  • Assign an owner to each affected ChatGPT, Work or Codex configuration scheduled for review before 14 October.
  • Check the current product documentation and actual account controls rather than relying on historical launch language.
  • Record the selected replacement available to that plan and workspace; do not infer a universal replacement from the API comparison.
  • Run the organisation’s required validation on the updated surface before treating the change as complete.

Direct API inventory checklist

  • Search application code, environment variables, deployment manifests, routing rules and test fixtures for gpt-5.5.
  • Distinguish aliases from fixed snapshots and record the exact identifier used in each environment.
  • Capture endpoint, reasoning effort, tool definitions, timeout policy, retry policy and output constraints.
  • Record representative input and output-token distributions, including the share of requests that exceed the documented 272,000-input long-context threshold.
  • Identify cache use, regional processing, tool charges, retries and human-review costs that prevent token-list prices from representing total workflow cost.
  • Verify current account, region, endpoint and rate-tier access independently of the statement that the retirement notice excludes the API.
  • Classify each deployment as “retain pending evidence”, “evaluate against GPT-5.4” or “change required for a separately documented reason”.

A compact decision scorecard for GPT-5.5 versus GPT-5.4

A scorecard should enforce acceptance requirements rather than convert every observation into one opaque average. A model that violates a mandatory safety, schema, tool or quality condition should not pass merely because it performs well on lower-risk cases.

Decision field What to record Gate or comparison?
Accepted quality Pass or fail under the workflow’s predeclared rubric, with error categories Gate first; compare passing configurations second
Safety and policy handling Observed handling of the representative risk cases selected by the team Gate where required by the workflow
Tool behaviour Tool selection, argument validity, stopping and recovery Gate for consequential actions; otherwise compare
Latency End-to-end task time, including tools and retries Compare against the service objective
Usage Input, cached-input and output tokens for the same cases Comparison input to cost calculation
Operational burden Retries, escalations, manual corrections and failed attempts Compare at workflow level
Total cost Cost per submitted task and per accepted result under declared assumptions Compare only after gates pass

Eliminate configurations that fail mandatory gates, then retain the lightest passing configuration.

Worked hypothetical examples

Example 1: A low-retry extraction workflow

Hypothetical example: A team evaluates 1,000 representative extraction records using fixed snapshots, identical prompts and the same output contract. Both models pass every mandatory schema and quality gate. GPT-5.4 also meets the latency objective and does not require materially more retries or human correction. Under those assumed results, the decision framework selects GPT-5.4 because it is the lighter passing configuration and has lower listed standard token rates.

This example does not establish that GPT-5.4 will match GPT-5.5 for extraction generally. It shows how the decision follows from a workload-specific result after both candidates pass the same acceptance contract.

Example 2: A tool workflow with fewer accepted outcomes

Hypothetical example: The same team tests a tool-enabled workflow. Suppose GPT-5.4 has a lower cost per attempt but requires enough retries and manual intervention that its cost per accepted result is higher under the team’s declared assumptions. Suppose GPT-5.5 passes the tool and stopping gates more often and its accepted-result cost falls within the approved budget. The team could justify GPT-5.5 for that workflow while continuing to use GPT-5.4 elsewhere.

The example illustrates why a two-to-one difference in listed standard token rates does not prove a two-to-one invoice or completed-task cost. The actual conclusion depends on token mix, caching, reasoning, tools, retries, long-context pricing conditions, processing choices and acceptance rates.

Example 3: A misleading default comparison

Hypothetical example: An initial test compares GPT-5.5 at its documented medium reasoning default with GPT-5.4 at its documented none default. GPT-5.5 produces stronger rubric scores but takes longer and uses more output tokens. The team should not label that result a model-only difference. It should rerun a configuration ladder, explicitly record reasoning.effort, and compare the lightest settings that meet the quality threshold.

If higher effort causes unnecessary tool exploration, fails to stop cleanly or performs worse under conflicting instructions, the correct response is not automatically to raise effort again. The team should inspect prompt conflicts, stopping conditions and tool scope before retesting.

Example 4: Surface retirement mistaken for API retirement

Hypothetical example: An engineering manager finds both a ChatGPT Work configuration and a backend service described internally as “GPT-5.5.” The Work configuration falls within the scheduled product-surface retirement review. The backend calls the direct API with gpt-5.5; the cited notice explicitly excludes the API. The manager opens separate change records: one for the Work update before 14 October and one for an API evaluation without inventing the same retirement deadline.

Failure-analysis procedure when results are ambiguous

  1. Confirm comparability. Verify that both runs used the same case set, prompt contract, tool definitions, endpoint conditions and scoring rules. Record snapshots and reasoning settings explicitly.
  2. Separate infrastructure failures from model outcomes. Timeouts, unavailable dependencies, malformed test fixtures and evaluator defects should not be counted as model-quality failures without classification.
  3. Cluster failures by type. Useful categories include instruction conflict, missing context, incorrect extraction, unsupported assertion, invalid tool arguments, unnecessary tool use, failure to stop, output-contract violation and reviewer disagreement.
  4. Inspect context placement and relevance. A listed 1,050,000-token context window establishes capacity, not reliable use of every supplied detail. Determine whether failures correlate with long, conflicting or poorly ordered inputs.
  5. Audit retry effects. Report first-attempt success separately from eventual success. A result reached after repeated calls has different latency and cost implications.
  6. Review reasoning-effort sensitivity. Test only justified settings. Higher effort can add latency and cost and, under the documented warning, can regress quality when instructions conflict, stopping is weak or tools are open-ended.
  7. Reconcile human judgments. Where reviewers disagree, resolve the rubric or label before attributing the disagreement to a model.
  8. Retest the corrected hypothesis. Change one material factor at a time where practical, preserve the prior run and document why the new run is comparable.
  9. Stop if evidence remains insufficient. An inconclusive evaluation supports neither a universal GPT-5.5 upgrade nor an automatic GPT-5.4 migration.

Failure report template

  • Case identifier and risk category
  • Model identifier, snapshot, endpoint and reasoning effort
  • Prompt and tool-contract version
  • Expected result and observed result
  • Failure category and reviewer evidence
  • Input, cached-input and output-token counts
  • Tool calls, retries and end-to-end latency
  • Whether the result was ultimately accepted
  • Proposed correction and retest status

Pre-release and rollback implementation checklist

Before changing a production route

  • Approve the representative case set and mandatory acceptance gates.
  • Pin or record snapshots where reproducibility is required.
  • Freeze the prompt, output contract, tools and runtime controls for the comparison.
  • Calculate cost using observed token mix and retry behaviour, not list-price ratios alone.
  • Include long-context, regional-processing and tool conditions where they apply.
  • Test rate-limit, timeout and recovery behaviour under the target account conditions.
  • Obtain the required safety, security and workflow-owner review.
  • Define the rollback trigger before release rather than after an incident.

Evidence required for rollback

  • The prior model identifier, snapshot and configuration remain recoverable.
  • Prompt and tool-contract versions can be restored together with the model route.
  • Monitoring distinguishes quality failures, tool failures, latency regressions and cost changes.
  • Rollback ownership and authorisation are named.
  • Requests affected during the change can be identified for review where the workflow requires it.
  • The decision log records whether rollback restores the full prior workflow or only the model identifier.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library →

Useful Links

  • OpenAI: Models — ChatGPT Work and Codex — current product documentation for the scheduled surface retirement and API exclusion; accessed 30 September 2026.
  • OpenAI: GPT-5.5 model reference — documented API specifications, pricing conditions, snapshot, reasoning defaults, endpoints and tools; accessed 30 September 2026.
  • OpenAI: GPT-5.4 model reference — documented API specifications, lower standard rates, snapshot, reasoning default, endpoints and tools; accessed 30 September 2026.
  • OpenAI: Using GPT-5.5 — guidance on representative baselines, reasoning-effort tuning and workflow controls; accessed 30 September 2026.
  • OpenAI: Model selection — guidance on same-input experiments and retaining the simplest passing configuration that meets the quality threshold; accessed 30 September 2026.
  • OpenAI: Introducing GPT-5.5 — the 23 April 2026 historical announcement and vendor-published launch evaluations, not current availability evidence or an independent benchmark.

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this