GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off

GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off

GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off

What OpenAI’s GPT-5.5 notice changes—and what it does not

OpenAI’s official ChatGPT notice says that GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex across all plans on October 14. The same notice directs Codex users to use GPT-5.6 Sol or GPT-6 Astra. Treat that as a product-surface retirement for the named ChatGPT, Work, and Codex experiences, not as proof of an API retirement. The notice, by itself, does not announce that GPT-5.5 is being retired from the OpenAI API; API availability should be verified against OpenAI’s API model documentation and any direct first-party API deprecation notice that applies to your account, project, or region.

This playbook is written for teams that cannot safely treat the October 14 change as a model-name swap. A migration from GPT-5.5 to GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-6 Astra through GPT-6 Pro, or another current model option can alter reasoning style, refusal boundaries, formatting, latency, cost exposure, tool-use behavior, and user expectations. The safe operating model is a controlled behavior-and-permission migration: inventory every dependency, capture current prompts and outputs, run representative evaluations, review tool permissions, compare costs and latency, define rollback, and require accountable human sign-off before production cutover.

OpenAI’s Help Center describes GPT-5.6 Sol as a paid-plan model for Instant and selectable thinking levels, GPT-5.6 Luna as the Free/Go default and Think model, GPT-5.6 Terra as a balanced option in Work, Codex, and the API, and GPT-6 Pro as powered by GPT-6 Astra on eligible plans. Those descriptions are a starting point for selection, not a guarantee that a specific workspace, account, app, or region exposes the same controls at the same time. Administrators should verify the actual model menu, workspace policy, allowance separation, and documentation visible to their own users before communicating a target state.

Codex migrations need extra care because the model is often coupled to repositories, shell commands, MCP servers, file edits, terminal approvals, and review workflows. OpenAI’s ChatGPT Work and Codex documentation says Work and Codex allowances are separate from Chat, and OpenAI’s GPT-5.6/GPT-6 Pro Help Center page states that Astra in Codex CLI requires version 0.153.0 or later. That means a model transition plan must include the client version, repository trust posture, permission profile, tool approvals, and the human review process—not only the destination model name.

Decision rule: classify the transition before you change anything

The first migration decision is whether your use of GPT-5.5 is casual, operational, or consequential. Casual use includes low-risk drafting, brainstorming, and personal productivity tasks where a different answer style is acceptable. Operational use includes reusable workflows, shared prompts, Work workspace automations, Codex code changes, document production, support triage, and internal analysis where output quality must be repeatable. Consequential use includes legal, financial, health, hiring, safety, youth, external publication, production code, customer communications, security operations, or any workflow where a bad output can create material harm. Only the first category can be handled with a lightweight user notice; the other two require a documented migration plan.

Migration class Typical examples Minimum control before cutover Human approval requirement
Casual productivity Brainstorming, personal summarization, non-sensitive draft outlines, learning exercises User communication, basic prompt adjustment, visible reminder that behavior can differ User accepts the changed model for their own work
Operational workflow Shared team prompts, Work workspace templates, Codex code-assistance workflows, support macros, recurring internal reports Inventory, representative fixtures, side-by-side review, cost and latency logging, permission review, rollback path Workflow owner signs off on acceptance criteria and known limitations
Consequential workflow External legal drafts, customer-facing messages, production code changes, security investigations, regulated-content workflows, financial or health-related analysis All operational controls plus privacy review, safety review, stricter failure thresholds, approval gates for external or destructive actions Accountable business, technical, and compliance owners approve before release

The conservative default is to classify ambiguous workflows as operational until evidence supports a lower-risk label. A prompt that “only drafts internal summaries” becomes consequential if those summaries influence employment, legal, medical, financial, or security decisions. A Codex session that “only proposes changes” becomes operational or consequential when it can edit files, run tests, call MCP tools, submit pull requests, alter infrastructure, or generate code that is merged into production.

Do not ask users to paste secrets, private customer records, privileged legal material, protected health information, payment data, authentication tokens, or restricted source material into migration prompts. If the evaluation requires sensitive inputs, use approved internal fixtures, synthetic substitutes, redacted samples, or a controlled environment governed by your organization’s privacy and security process. Human approval remains mandatory for external messages, submissions, payments, purchases, bookings, destructive actions, permission changes, publication, legal commitments, campaign launches, and any other consequential operation.

The operating model: behavior, permissions, evidence, and sign-off

A reliable GPT-5.5 migration has four workstreams running in parallel. The first is behavior: collect representative tasks, expected formats, evaluation criteria, and known failure modes. The second is permissions: identify every tool, repository, workspace connector, MCP server, file path, browser action, and command path that the assistant can touch. The third is evidence: run side-by-side comparisons with logs for output quality, refusal behavior, cost, latency, and tool-use outcomes. The fourth is sign-off: require named owners to accept the results, risks, rollback procedure, and user communication plan.

OpenAI’s evals guidance supports the central idea that teams should test models against representative examples rather than rely on a single manual trial. For this migration, representative means the fixture set reflects the tasks users actually perform with GPT-5.5: short and long prompts, easy and hard cases, common and edge-case file types, normal and adversarial instructions, allowed and disallowed tool calls, expected refusals, and examples that previously required careful human correction. A clean evaluation set prevents the migration from being judged only by impressive demos or by a single failure that does not match real usage.

OpenAI’s prompt-engineering guidance should be applied as a change-control discipline, not as a last-minute rewrite session. Store system instructions, developer instructions, reusable user prompts, output schemas, tool descriptions, and evaluation rubrics under version control or another approved change-tracking system. The migration team should be able to answer which prompt version was tested, which model was used, which tool permissions were enabled, what outputs were accepted, what failures were observed, and who approved the result.

For Codex, the same operating model must include repository state. Before asking Codex to modify a codebase, establish a Git checkpoint, confirm the branch, capture dependency versions, define the test command, and state which files or directories are in scope. OpenAI’s Codex CLI documentation recommends Git checkpoints before and after tasks, and that practice becomes more important during a model transition because it gives reviewers a precise diff to inspect and a known rollback point if the new model behaves differently.

Opening migration checklist for October 14 readiness

Use the following checklist before any workspace-wide announcement says “we migrated.” The checklist is intentionally front-loaded because most migration failures are discovered too late: a forgotten shared prompt, a Codex permission profile that is broader than intended, a report template that depends on a prior formatting habit, or a cost pattern that only appears when dozens of users repeat the same task. The goal is not to eliminate all differences between GPT-5.5 and the replacement model; the goal is to make the differences visible, bounded, approved, and reversible where possible.

  1. Confirm scope from first-party sources. Record that OpenAI’s ChatGPT notice retires GPT-5.5 from ChatGPT, ChatGPT Work, and Codex across all plans on October 14, and that Codex users are directed to GPT-5.6 Sol or GPT-6 Astra. Separately check the API model documentation before making any API retirement claim.
  2. Name the migration owner. Assign one accountable owner for each workspace, product, repository group, department, or workflow family. The owner must have authority to pause cutover if evaluation evidence is insufficient.
  3. Inventory GPT-5.5 dependencies. Capture saved prompts, custom instructions, Work templates, Codex workflows, repository tasks, MCP tools, knowledge files, document workflows, browser or desktop actions, and user-facing outputs that depend on GPT-5.5 behavior.
  4. Classify risk. Label each workflow as casual, operational, or consequential. Escalate anything involving external publication, code changes, legal or compliance review, security operations, personal data, regulated content, or customer commitments.
  5. Select candidate replacement models. Use current first-party documentation and your actual account availability to choose candidates such as GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, or GPT-6 Astra through eligible product surfaces. Do not promise a model to users until availability is verified in the relevant surface.
  6. Freeze representative fixtures. Build a test set from real workflow patterns, using synthetic or redacted data when sensitive material is not appropriate. Include expected output contracts and known unacceptable outputs.
  7. Review tool permissions. For Codex and tool-enabled workflows, document allowed tools, disabled tools, approval requirements, sandbox boundaries, network assumptions, repository scope, and any MCP server configuration.
  8. Run side-by-side tests. Compare GPT-5.5 outputs where still available against the candidate model outputs. Record pass, fail, partial, latency, cost indicators where visible, reviewer notes, and required prompt changes.
  9. Set rollback rules. Define what will be reverted, paused, disabled, or rerouted if the replacement model fails acceptance thresholds. Rollback can include restoring prior prompts, disabling a tool, reverting a repository change, pausing a shared template, or routing tasks to human-only review.
  10. Obtain sign-off. Require named approval from the workflow owner and, where applicable, security, legal, compliance, engineering, support, education, or business stakeholders before the migrated workflow becomes the default.

Why model-name swaps fail in real organizations

A model-name swap fails when the organization assumes that the surrounding workflow is stable. In practice, the model is only one component in a chain that includes user intent, system instructions, workspace policy, model availability, reasoning settings, tool permissions, file access, retrieval context, output schemas, review steps, and downstream systems. If any part of that chain changes, a previously acceptable prompt can produce a different format, a different level of detail, a different refusal, a different tool call, or a different interpretation of the same source material.

For knowledge workers, the common failure is silent format drift. A GPT-5.5 prompt that reliably produced a five-bullet executive summary may produce a longer narrative, different caveats, or a different prioritization under a newer model. That is not automatically worse, but it can break a recurring meeting workflow, a CRM note convention, a classroom handout, or a legal work-product review checklist. The migration fixture should therefore include expected structure, not only a vague instruction to “summarize well.”

For developers and Codex users, the common failure is permission mismatch. A new model may reason differently about when to inspect files, run commands, propose refactors, or ask for clarification. If the permission profile is too broad, reviewers may face larger diffs or unexpected command requests; if it is too narrow, the model may stall or produce speculative code without enough context. The solution is not to disable approvals or grant broad access by default. The solution is to state the task boundary, keep approvals in place, use Git checkpoints, and evaluate whether the model’s tool-use pattern is acceptable for the repository.

For enterprise administrators, the common failure is communicating a single replacement model before confirming plan and workspace constraints. OpenAI’s documentation describes model roles and product surfaces, but actual access can vary by plan, account, region, rollout, and workspace policy. A migration memo that says “everyone should use GPT-6 Astra” can create confusion if only eligible users see GPT-6 Pro, if Codex client versions differ, or if Work permissions expose a different model set. Administrators should communicate verified options per surface, not broad assumptions.

For legal-technology professionals, the common failure is over-reading customer stories or model names as quality guarantees. OpenAI has described GPT-6 Astra in legal-document contexts through customer stories, but those observations are not independent legal-quality certifications and do not remove qualified lawyer review. A legal workflow migration should test citation handling, formatting, issue spotting, source grounding, privilege boundaries, redaction, and review annotations against local standards before any external or client-facing use.

For educators and parents, the common failure is treating a model migration as a purely technical event. If students, minors, or family members use ChatGPT for tutoring, writing feedback, or coding help, the change may affect tone, scaffolding, refusal behavior, and how much reasoning is shown or requested. Schools and families should avoid entering sensitive student records or personal identifiers into ad hoc tests, should verify age-appropriate settings and institutional policies, and should keep human review for grading, discipline, accommodations, mental-health concerns, or safety-related decisions.

Source-grounded model-selection frame for the first pass

OpenAI’s Help Center positions GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.6 Terra, and GPT-6 Pro powered by GPT-6 Astra across different product contexts. For this migration, use those descriptions to form a candidate shortlist, then validate against your own tasks. A paid-plan ChatGPT workflow might begin with GPT-5.6 Sol if it needs Instant and selectable thinking levels as documented by OpenAI. A Work or Codex workflow may consider GPT-5.6 Terra where it is available as a balanced option. A Codex team directed by the ChatGPT notice may test GPT-5.6 Sol and GPT-6 Astra, while also confirming Codex CLI version requirements for Astra where relevant.

The first pass should not ask “which model is best?” because that question is too broad to operationalize. Ask “which available model passes this workflow’s acceptance criteria with acceptable permissions, cost, latency, and review burden?” A model that is excellent for deep reasoning may be a poor default for high-volume short rewrites if it increases latency or review friction. A balanced model may be preferable for routine internal transformations but insufficient for complex codebase reasoning. A more advanced model may be justified for high-risk drafting only if it reduces material review failures under your rubric.

Selection question Evidence to collect Decision rule
Does the candidate preserve required output structure? Side-by-side outputs against fixed schemas, reviewer notes, examples of format drift Accept only if downstream users or systems can consume the output without unsafe manual repair
Does the candidate handle hard cases better, worse, or differently? Representative edge cases, known prior failures, adversarial or ambiguous prompts where appropriate Accept only if differences are understood and documented, not merely surprising in a demo
Does tool use remain bounded? Requested commands, file edits, MCP calls, browser actions, approval prompts, denied attempts Accept only if permissions match the task and consequential actions remain behind human approval
Is the review burden acceptable? Number of corrections, severity of errors, reviewer time, unresolved uncertainty Accept only if human reviewers can reliably catch and correct remaining failures before use
Are cost and latency acceptable? Observed response time, usage indicators available in the product or API context, retry frequency Accept only if the workflow owner agrees the benefit justifies the operational cost

Minimum evidence package before teams begin broader testing

Before inviting a wider group of users to test the replacement model, assemble a minimum evidence package. This package prevents the migration from becoming a collection of anecdotes and gives security, compliance, engineering, and business owners a shared basis for approval. It also protects users from contradictory instructions, such as one team asking for broad tool access while another team has not reviewed the same permission profile.

  • Scope statement: The exact product surfaces covered, such as ChatGPT, ChatGPT Work, Codex, or API-backed applications. The statement must explicitly note that OpenAI’s ChatGPT notice does not itself announce GPT-5.5 API retirement.
  • Dependency inventory: A list of prompts, templates, Codex tasks, repositories, documents, connectors, MCP servers, tool permissions, and user groups affected by the change.
  • Model shortlist: Candidate models selected from currently documented and actually available options for the relevant plan, workspace, app, and region.
  • Fixture set: Representative examples with redacted or synthetic data where needed, expected outputs, scoring criteria, and explicit disallowed outcomes.
  • Prompt versions: Captured instructions, prompt changes, output schemas, and reviewer rubrics with version identifiers.
  • Permission matrix: Allowed tools, denied tools, required approvals, sandbox assumptions, repository boundaries, and escalation rules for uncertain tool requests.
  • Comparison log: Results from old and candidate models where available, including quality observations, formatting differences, refusals, hallucination risks, tool-use traces, latency, and cost indicators.
  • Risk register: Known unresolved differences, mitigations, owner names, and conditions that would block or delay cutover.
  • Rollback plan: Steps to restore prior prompts, disable risky workflows, revert code changes, pause shared templates, or route tasks to manual review.
  • Sign-off record: Named approval from accountable owners and, where appropriate, security, compliance, legal, education, support, or engineering leaders.

The evidence package should be lightweight enough to complete before October 14 but structured enough to survive audit, incident review, or leadership questions. A ten-row spreadsheet with clear fixtures and sign-off can be better than a long narrative that does not identify owners, thresholds, or rollback actions. The practical test is whether a new reviewer can understand what changed, why it was accepted, what risks remain, and how to stop the workflow if those risks materialize.

Sample migration charter for the project kickoff

The following sample charter is a recommendation, not an OpenAI template. Adapt it to your organization’s change-management system and do not include credentials, private user data, customer secrets, privileged legal material, or regulated records. Its purpose is to align the migration team before anyone rewrites prompts, changes Codex defaults, or tells users that a replacement model is approved.

Migration charter: GPT-5.5 product-surface transition

Official source basis:
- OpenAI's ChatGPT notice states that GPT-5.5 retires from ChatGPT,
  ChatGPT Work, and Codex across all plans on October 14.
- The notice directs Codex users to GPT-5.6 Sol or GPT-6 Astra.
- This charter does not treat the notice as an API retirement announcement.
  API availability must be checked against OpenAI API model documentation
  and any first-party API deprecation notice applicable to our account.

Scope:
- Product surfaces in scope: [ChatGPT / ChatGPT Work / Codex / API application name]
- Workflows in scope: [workflow names]
- User groups in scope: [teams or roles]
- Repositories or systems in scope: [authorized systems only]

Candidate models:
- Candidate 1: [model name visible and permitted in this product surface]
- Candidate 2: [model name visible and permitted in this product surface]
- Selection constraints: [plan, workspace policy, region, client version, permissions]

Evaluation:
- Fixture set owner: [name or role]
- Number and type of representative tasks: [description]
- Sensitive-data handling: [synthetic / redacted / approved controlled data]
- Acceptance criteria: [quality, format, refusal behavior, tool-use, cost, latency]
- Blockers: [conditions that prevent cutover]

Permissions:
- Tools allowed: [specific tools]
- Tools denied: [specific tools]
- Approval required for: external messages, submissions, payments, purchases,
  bookings, destructive actions, permission changes, publication, legal commitments,
  campaign launches, production code changes, and other consequential operations.
- Codex repository boundaries: [paths, branches, review rules]

Rollback:
- Prompt rollback location: [version-controlled location or approved system]
- Code rollback method: [Git revert / branch reset / release rollback procedure]
- Workflow pause owner: [name or role]
- User communication owner: [name or role]

Sign-off:
- Business owner: [name]
- Technical owner: [name]
- Security/compliance/legal owner if applicable: [name]
- Final approval date: [date]

A charter like this is especially useful when different groups control different parts of the migration. Enterprise administrators may control workspace policy, developers may control Codex configuration, security may control connector approval, legal may control privileged workflows, and business teams may own output quality. Without a charter, each group can believe another group validated the risk that actually sits between them.

Immediate next step: build the dependency inventory before evaluating models

The next section of the playbook should begin with inventory because evaluation without inventory produces false confidence. If a team tests three favorite prompts and ignores the shared prompt library, Codex repositories, Work workspace templates, and user-created procedures, it may approve a migration that only covers the easiest visible cases. The inventory is also where you discover permission risks: a workflow that looked like text generation may actually depend on file uploads, browsing, MCP tools, code edits, or external publication.

Start the inventory by asking every workflow owner to list where GPT-5.5 is deliberately selected, where it is implicitly selected by an old template, and where users may have copied a GPT-5.5-specific prompt into personal notes. Include prompt text, expected output, model setting, reasoning setting if visible, tool access, data sensitivity, downstream consumer, review owner, and fallback path. Do not collect real secrets or unnecessary confidential content; record the existence of sensitive data categories rather than copying the data into the migration tracker.

The operational standard for the rest of this playbook is simple: no production cutover without representative evidence, no tool expansion without permission review, no consequential action without human approval, no API claim without first-party API documentation, and no user communication that promises behavioral equivalence. That standard may feel slower than changing a model selector, but it is faster than discovering after October 14 that the organization migrated the name while leaving behavior, permissions, cost, and accountability untested.

Build the inventory before judging the replacement model

GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off — first editorial explainer visual

The migration inventory is the control plane for the rest of the GPT-5.5 transition. OpenAI’s official ChatGPT account says GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex across all plans on October 14 and recommends GPT-5.6 Sol or GPT-6 Astra for Codex users; that notice does not, by itself, establish an API retirement schedule. Treat product surfaces and API integrations as separate scopes until a first-party API source says otherwise. The inventory should therefore record where GPT-5.5 is selected in ChatGPT, ChatGPT Work, Codex, and any API-facing configuration, while preserving the distinction between a product migration and an API migration.

A usable inventory is not a list of “places we think people use ChatGPT.” It is a versioned register of prompts, saved instructions, workspace policies, repositories, tools, model settings, output contracts, data classes, human approvers, usage allowances, and rollback owners. OpenAI’s Help Center states that Work and Codex allowances are separate from Chat, and that current model options vary by plan and product; your inventory must capture allowance boundaries instead of assuming that a successful Chat migration proves Codex capacity or Work availability.

Start with read-only discovery. Do not ask employees to paste secrets, confidential source code, customer records, legal files, health information, or privileged material into a migration spreadsheet or chat. Ask owners to provide metadata: product surface, workspace, repository path, task category, model currently selected, tools enabled, data classification, output destination, reviewer role, failure impact, and the location of version-controlled prompt artifacts. If confidential examples are needed for evaluation, use approved test fixtures inside the appropriate environment and record only redacted evidence in central reports.

Inventory object What to capture Why it matters during migration Owner to name
Chat conversations and saved instructions Task family, saved instruction summary, model preference, reasoning setting if selectable, file/input types, output format, reviewer role Saved instructions often contain implicit style rules, forbidden outputs, or workflow assumptions that change behavior when the model changes Business process owner or knowledge-work lead
ChatGPT Work spaces Workspace, team, allowed models, shared prompts, data classifications, approval requirements, usage allowance dependencies Work availability and allowances are separate from Chat, so capacity and permission testing must be scoped to the workspace Workspace administrator and department approver
Codex repositories Repository, branch policy, trusted project status, Codex version, model selection, tool permissions, sandbox assumptions, test command set Codex migrations can change code-generation behavior, review quality, terminal-command proposals, and repository-specific automation paths Engineering manager, repository maintainer, or platform owner
Prompts and prompt chains Prompt ID, version, system/developer/user prompt text location, variables, expected output schema, disallowed content, sample inputs Prompt changes and model changes must be separated; otherwise teams cannot tell whether the model or prompt caused a regression Prompt maintainer or product owner
Skills, plugins, tools, and connectors Tool name, purpose, approval mode, scopes, enabled users, external systems touched, rate/cost impact, failure modes Model replacement can alter when tools are invoked, what arguments are drafted, and how confidently the assistant explains tool results Tool owner, security owner, and data owner
Output schemas and templates JSON schema, document template, citation format, code style, file naming, validation command, downstream parser Downstream systems can fail even when prose looks better, especially where parsers expect exact keys, order, or typed fields Application owner or workflow integrator
Data classes Public, internal, confidential, regulated, privileged, youth-related, security-sensitive, or customer-provided classification Migration testing must not widen access to sensitive data or move restricted data into unauthorized prompts, logs, or review tools Data steward, legal, privacy, or security representative
Human approvers Required reviewer role, backup reviewer, approval deadline, escalation path, sign-off artifact Consequential outputs require accountable human acceptance before rollout, publication, submission, purchase, code merge, or customer delivery Process owner and risk owner

Inventory fields that prevent ambiguous migration reports

Every entry should have a stable identifier. A practical format is SURFACE-TEAM-WORKFLOW-SEQUENCE, such as CODEX-PLATFORM-REVIEW-017 or WORK-LEGAL-SUMMARY-004. Stable IDs let reviewers compare GPT-5.5 baseline outputs, GPT-5.6 candidate outputs, Codex task traces, prompt versions, tool settings, cost notes, and sign-off records without relying on conversation titles or screenshots.

Record the product surface as one of at least four categories: ChatGPT personal or individual use, ChatGPT Work, Codex, and API. If your organization uses additional internal wrappers, classify them separately and map them back to the official surface that provides the model. This prevents the common mistake of applying a ChatGPT retirement notice to an API integration, or assuming that an API model catalog entry means the same behavior, controls, or availability exists in ChatGPT Work or Codex.

Capture the current and proposed model as an evidence field, not a wish list. OpenAI’s Help Center describes GPT-5.6 Sol as available for paid-plan Instant and selectable thinking levels, GPT-5.6 Luna as the Free/Go default and Think model, GPT-5.6 Terra as a balanced option in Work, Codex, and the API, and GPT-6 Pro as powered by GPT-6 Astra on eligible plans. Because eligibility varies, inventory records should say “candidate model to test” and “confirmed available in this workspace” as separate fields.

Reasoning controls deserve their own column when the product surface exposes them. A model name alone is insufficient if one workflow uses an instant setting, another uses a thinking mode, and another relies on Codex behavior. The evaluation fixture should preserve the exact setting used in the baseline where possible and clearly document the replacement setting when exact parity is unavailable. Do not infer hidden reasoning behavior; record only the selectable setting, user-facing option, or API parameter that your team actually controls.

Usage allowance and budget exposure should be captured before evaluation begins. OpenAI’s Help Center states that Work and Codex allowances are separate from Chat, so high-volume Codex trials can fail operationally even when Chat trials pass. Inventory each workflow’s expected run count, typical input size, output size class, tool activity, and review burden. Use this to prioritize small canaries before broader shadow runs, and to identify workflows that need administrator confirmation of allowance, spend, or rate constraints.

{
  "workflow_id": "CODEX-PLATFORM-REVIEW-017",
  "surface": "Codex",
  "owner": "Platform Engineering",
  "current_model_recorded": "GPT-5.5",
  "candidate_models_to_test": ["GPT-5.6 Sol", "GPT-5.6 Terra", "GPT-6 Astra if eligible"],
  "availability_confirmed_by": "workspace_admin_name_or_role",
  "prompt_artifact": "repo/path/prompts/review_prompt_v3.md",
  "output_contract": "repo/path/schemas/review_findings.schema.json",
  "tools_enabled": ["repository_read", "test_runner"],
  "tools_requiring_human_approval": ["terminal_command", "pull_request_creation"],
  "data_class": "internal_source_code",
  "representative_fixtures": ["fixture_001", "fixture_014", "fixture_022"],
  "critical_error_classes": ["unsafe_command", "missed_security_regression", "fabricated_test_result"],
  "rollback_owner": "platform_release_manager",
  "human_signoff_required": true
}

This JSON is an example record structure, not an OpenAI endpoint or required schema. Keep real secrets, access tokens, private keys, customer identifiers, privileged legal facts, and sensitive personal data out of central migration records unless your approved governance process explicitly requires and protects them. The central register should usually point to controlled repositories, ticket systems, and evaluation artifacts rather than duplicating sensitive content.

Saved instructions, custom prompts, and hidden workflow assumptions

Saved instructions can be more important than the prompt body. A finance analyst may have a saved instruction requiring conservative assumptions and a specific spreadsheet explanation style. A support lead may have a standing instruction to avoid refund promises. A legal-technology user may have instructions about jurisdictional disclaimers and source hierarchy. During migration, capture these as controlled prompt dependencies, because replacing the model without capturing the instruction context can produce misleading evaluation results.

For each saved instruction or reusable prompt, separate four layers: purpose, policy constraints, style preferences, and output contract. Purpose describes the job, such as “summarize customer escalation notes for internal triage.” Policy constraints define prohibitions, such as “do not infer medical status” or “do not provide legal advice.” Style preferences define tone, length, headings, or reading level. Output contract defines exact sections, fields, citations, JSON keys, or table columns. This separation helps reviewers decide whether a changed answer is an acceptable style shift or a functional regression.

Prompt version control should include the model-transition context. Store prompts in a repository or governed document system with a version identifier, owner, approval history, and changelog. A minimum naming pattern is prompt-name.major.minor.patch, where major changes alter task behavior, minor changes alter instructions or examples, and patch changes correct wording without changing requirements. Do not silently edit prompts while comparing models; freeze the prompt version for the first side-by-side run, then run a separate prompt-tuning phase only after baseline deltas are understood.

Output schemas need the same treatment. If a workflow expects JSON, define required keys, allowed values, nullable fields, maximum length constraints where relevant, and validation commands. If a workflow expects prose, define the required sections, citation behavior, refusal behavior, and disallowed claims. For code, define formatting, linting, tests, security checks, and review criteria. Model migrations fail quietly when teams accept “looks better” while breaking a downstream parser, checklist, or reviewer expectation.

Repository, Codex, and tool-permission inventory

Codex entries require a more operational inventory than ordinary chat tasks because repository access, terminal commands, MCP tools, sandbox rules, and approval policies can affect the outcome. OpenAI’s ChatGPT Work and Codex documentation says Work and Codex are distinct product surfaces, and the broader Codex documentation describes model selection and permissions as explicit controls. For migration purposes, record the Codex version, repository trust status, project configuration location, enabled tools, disabled tools, network expectations, and which actions require review before execution.

Do not treat tool permissions as a static copy of the old setup. A different model can be more or less likely to request a tool, generate a command, change files, or propose a pull request. The migration review should ask whether each tool is still necessary, whether the approval mode is still appropriate, whether scopes are narrow enough, and whether a human must approve before external messages, submissions, purchases, bookings, deployments, permission changes, destructive file operations, or publication. A model change is a good forcing function for least-privilege review.

For repositories, inventory the task types rather than only the codebase name. “Uses Codex” is too broad. Break it into code review, test generation, bug triage, refactoring, documentation edits, migration scripts, dependency updates, incident analysis, and release note drafting. Each task type has different acceptance criteria and rollback needs. A documentation edit may be reversible through Git; a migration script that touches production data requires stricter approval, dry-run evidence, backups, and a qualified owner.

Record the commands that evaluators may run and the commands that are prohibited. Approved examples might include read-only test discovery, unit tests in a sandbox, format checks, or schema validation. Prohibited examples should include commands that alter production systems, expose credentials, disable security controls, erase data, change permissions, send external communications, or bypass review. If elevated terminal input approval appears in the workflow, require explicit human confirmation and evidence that the command is necessary for the fixture.

Permission area Inventory question Migration decision rule
Repository write access Can the assistant modify files, open branches, or draft pull requests? Allow only inside authorized repositories with Git checkpoints and human review before merge or publication.
Terminal execution Which commands are allowed, which require approval, and which are never allowed? Prefer read-only discovery and sandboxed tests; require human approval for elevated, destructive, networked, or state-changing commands.
External tools and connectors Which systems can receive data or actions from the model-assisted workflow? Use least privilege, narrow scopes, and approval gates for consequential actions or data movement.
MCP or plugin-like tools Which tool names, scopes, authentication modes, and approval settings are enabled? Reconfirm necessity, callback/configuration ownership, and per-tool approvals before migration testing.
Network access Does the workflow need internet, internal services, package registries, or no network access? Default to the narrowest access that can complete the representative fixture; do not broaden network permissions to make a migration pass.

Define representative fixtures before running side-by-side comparisons

Representative fixtures are the test cases that decide whether the replacement is acceptable. OpenAI’s Evals guidance supports structured evaluation of model behavior, and OpenAI’s prompt-engineering guidance emphasizes clear instructions and examples. For this migration, fixtures should be drawn from real task families without exposing unnecessary sensitive data. A fixture must include the input, permitted context, prompt version, model setting, tool permissions, expected output contract, scoring method, critical error classes, and review owner.

Use stratified sampling rather than convenience sampling. Include common easy cases, high-volume routine cases, edge cases, ambiguous cases, historically problematic cases, and high-risk cases where a mistake would have legal, security, financial, safety, reputational, or operational consequences. If a team only tests clean examples, the migration report will overstate readiness. If a team only tests pathological cases, it may reject a model that is suitable for routine work but needs routing rules for special cases.

A good fixture is reproducible. It should specify the exact prompt artifact, allowed files, redacted input package, expected output type, reviewer rubric, and any tools that may be used. For Codex, include the repository commit or checkpoint, branch setup, test command, and prohibited operations. For ChatGPT Work, include workspace assumptions and whether uploaded files are synthetic, redacted, or approved internal material. For API-adjacent testing, include the request shape only if it is in scope and authorized; do not infer product-surface behavior from API-only tests.

Define the baseline before seeing the candidate output. For each fixture, preserve the GPT-5.5 output where available, the human-accepted final output if one exists, and any known defects in that baseline. The goal is not to worship GPT-5.5 behavior; it is to understand what users depended on. A replacement can be better and still fail migration if it breaks a required schema, omits a mandatory warning, overuses tools, fabricates sources, or changes a decision workflow without approval.

Fixture ID: WORK-COMPLIANCE-SUMMARY-006
Surface: ChatGPT Work
Task family: Internal policy summarization
Data class: Internal, non-public, no personal data in fixture
Prompt version: compliance-summary-v2.1.0
Candidate model setting: GPT-5.6 option confirmed available in workspace
Input package: redacted policy excerpt and reviewer question set
Expected output:
  - One-paragraph executive summary
  - Three risk bullets
  - "Unknowns and required human review" section
  - No legal advice phrasing
Critical error classes:
  - Invents policy requirements not present in the source
  - Omits required human-review section
  - Recommends external publication or enforcement without approval
Scoring:
  - Schema pass/fail
  - Factual support score by blind reviewer
  - Safety/policy review pass/fail
Acceptance threshold:
  - Zero critical errors
  - At least 4 of 5 on factual support for routine rollout
  - Human owner sign-off required before use

This fixture format is intentionally plain. It can live in a repository, governance document, or evaluation tool. The important property is that reviewers can rerun the same case, inspect the same input boundaries, and determine whether a candidate output passed for the right reasons.

Acceptance thresholds and critical error classes

Acceptance thresholds must be defined before testing begins. If the threshold is invented after reviewers see the outputs, the migration becomes a preference debate. A practical threshold has three layers: absolute blockers, scored quality gates, and operational gates. Absolute blockers include critical safety, security, privacy, legal, or destructive-action errors. Scored quality gates measure usefulness, factuality, completeness, schema compliance, and maintainability. Operational gates measure latency, allowance consumption, review time, tool activity, and rollback readiness.

Critical error classes should be tailored to the workflow, but several classes recur across Chat, Work, and Codex. A hallucinated citation in a legal-technology summary, a fabricated test result in a Codex review, an unsafe command proposal, a privacy leak in a customer-support draft, a schema-breaking JSON response, or a missing human-approval warning can each be a release blocker. Treat critical errors as binary blockers unless a named risk owner approves a narrower rollout with compensating controls.

Critical error class Example in migration testing Default consequence
Unsupported factual claim The output invents a policy, source, case citation, test result, customer fact, or product capability not present in the input or official source. Block the fixture; require prompt, retrieval, source, or reviewer-control changes before rollout.
Privacy or confidentiality breach The workflow moves restricted data into an unauthorized prompt, log, tool, repository, or shared report. Stop testing for that workflow and escalate to privacy/security owners.
Unsafe or unauthorized action The assistant proposes or executes a destructive command, permission change, external submission, payment, deployment, or message without approval. Block rollout and tighten tool permissions, approvals, and instructions.
Output contract failure Required JSON keys are missing, a legal disclaimer section is omitted, or a downstream parser fails. Block automated or semi-automated use until the schema passes consistently.
Review deception or overconfidence The output claims tests passed when they were not run, hides uncertainty, or presents speculative conclusions as verified. Block high-risk use and require explicit evidence reporting in the prompt contract.
Cost or latency outlier A routine workflow consumes materially more allowance or reviewer time than the baseline budget permits. Route to cost review; consider different model, prompt compression, fixture routing, or narrower use.

For scored criteria, use a small rubric that reviewers can apply consistently. A five-point factual-support score is useful only if “5” means all material claims are supported by the provided context, “3” means minor unsupported wording without decision impact, and “1” means material unsupported claims. For code, use objective gates first: tests pass, lint passes, diff is scoped, no prohibited files changed, and the explanation matches the diff. Subjective readability can be secondary.

Do not require the candidate model to imitate every GPT-5.5 phrasing choice. Behavioral equivalence is not guaranteed, and exact imitation may be the wrong goal. Require preservation of obligations: correct facts, correct permissions, correct schema, correct human-review boundaries, acceptable cost, and acceptable user experience. If users prefer the old tone, handle that through prompt iteration after the first evaluation, not by hiding functional regressions under style debates.

Blind review and side-by-side comparison rules

Blind review reduces model-preference bias. When possible, show reviewers anonymized outputs labeled “A” and “B” without revealing the model. Ask reviewers to score against the fixture rubric before discussing preferences. If a reviewer knows the model because the interface exposes it, record that limitation. The goal is not academic purity; it is to avoid choosing the newer model because it sounds more polished or rejecting it because it differs from familiar GPT-5.5 phrasing.

Use at least two reviewer roles for high-risk workflows: a domain reviewer and an operational reviewer. The domain reviewer checks facts, policy, legal-adjacent reasoning, educational appropriateness, safety language, or business correctness. The operational reviewer checks schema, tool use, permissions, latency, cost, evidence capture, and rollback notes. For Codex tasks, include a maintainer who can inspect diffs and a security or platform reviewer for commands, dependencies, authentication, and deployment impact.

Side-by-side review should compare four artifacts, not two answers. Review the baseline prompt and output, the candidate prompt and output, the tool trace or action log where available, and the human final answer or merged diff when one exists. This prevents a candidate from being penalized for correcting a GPT-5.5 weakness, and it prevents a fluent answer from passing when it skipped a required external check, test, or approval gate.

Write disagreement rules before reviewers start. If reviewers disagree on a non-critical score, use a third reviewer or average score only when the rubric allows it. If any reviewer identifies a critical error, the fixture should fail pending triage. If a workflow owner wants to accept a known weakness, require a documented exception with scope, duration, compensating controls, rollback plan, and named accountable approver.

Separate product migration scope from API evaluation scope

The official retirement notice cited for this playbook concerns ChatGPT, ChatGPT Work, and Codex product surfaces. OpenAI’s API model documentation is a separate source for API model availability. Unless an official API source states that a particular GPT-5.5 API model is retiring on the same schedule, do not tell engineering teams that their API integration is automatically subject to the October 14 product-surface retirement. Conversely, do not assume an API model catalog entry proves that a ChatGPT Work or Codex workspace can select the same model with the same controls.

This distinction changes the evaluation contract. Product-surface evaluations should test the actual ChatGPT, Work, or Codex experience: saved instructions, workspace policies, selectable model options, reasoning settings, file handling, tool approvals, and user review flows. API evaluations should test request parameters, schemas, streaming behavior if used, error handling, latency, costs, rate limits, and application-level safety controls. A pass in one scope can inform the other, but it cannot substitute for it.

For founders and product leads, the decision rule is simple: if users interact with the ChatGPT or Codex interface, run product-surface fixtures; if your application calls OpenAI through code, run API fixtures against the official API model documentation and your production-like staging environment. If both are true, maintain two sign-off tracks. This avoids the common failure mode where a workspace migration looks complete while a service account, wrapper, or background job still depends on a separate model selection.

Output version control and evidence retention

Keep the evidence package lightweight but durable. Each workflow should retain the fixture ID, prompt version, candidate model, product surface, selectable reasoning setting if any, date, reviewer identities or roles, scores, critical errors, cost or allowance notes, latency notes if measured, tool actions, and final decision. Retain redacted outputs where policy permits. If sensitive content cannot be centrally stored, retain a pointer to the approved evidence location and a redacted decision summary.

Use immutable or append-only records for sign-off decisions where possible. Migration reviews become unreliable when teams overwrite old outputs or edit prompts without a changelog. A simple repository structure can work: /fixtures for inputs, /prompts for prompt versions, /schemas for output contracts, /runs for redacted run records, and /decisions for approval notes. For regulated, privileged, or confidential workflows, use your organization’s approved records system instead of a general repository.

Cost and latency evidence should be collected at the same time as quality evidence, not after rollout. A candidate output that is better but too slow for a live support queue, too expensive for a high-volume enrichment job, or too demanding of Work or Codex allowances may need routing rules. Record representative run duration, retry count, reviewer time, tool invocations, and allowance impact where your environment exposes those measures. Do not invent costs or limits; use the current official documentation and your own administrator-visible usage data.

Decision record template:

Workflow ID:
Surface:
Owner:
Prompt version:
Output contract version:
Candidate model and setting:
Fixture set:
Number of runs:
Critical errors observed:
Quality threshold result:
Schema/tool threshold result:
Cost/latency or allowance notes:
Privacy/security review result:
Required compensating controls:
Rollback owner:
Decision: approve / approve limited canary / reject / retest required
Approvers:
Date:
Evidence location:

The decision record is a management artifact, not a substitute for human judgment. Human approval remains mandatory before external messages, customer-facing outputs, code merges, production deployments, legal submissions, financial commitments, purchases, bookings, permission changes, publication, or other consequential operations. The model can assist the evaluation; it should not be the final authority on its own readiness.

Minimum evaluation contract before model testing begins

Before any team declares a replacement model acceptable, require a written evaluation contract. The contract should state the product surface, candidate models, fixtures, prompt versions, output schemas, tools, data classes, review roles, thresholds, critical error classes, and rollback rule. This contract prevents the migration from turning into ad hoc experimentation where every team changes prompts, models, tools, and acceptance criteria at the same time.

The contract should also state what is out of scope. If the task is a ChatGPT Work policy-summary migration, the contract should not silently approve Codex repository changes. If the task is Codex code review, it should not approve production deployment or dependency upgrades. If the task is API evaluation, it should not imply that ChatGPT Work users can access the same model. Explicit exclusions are useful because model migrations often expand as teams discover adjacent workflows.

Recommended migration rule: no workflow moves from GPT-5.5 dependence to a replacement model on trust alone. It needs an inventory record, representative fixtures, frozen prompt and output versions, permission review, acceptance thresholds, blind or structured review where feasible, cost or allowance notes, rollback ownership, and named human sign-off.

A compact contract can fit on one page for low-risk workflows and may require a full test plan for high-risk workflows. The form is less important than the discipline: evaluate the work users actually do, on the surface where they actually do it, with the permissions and review gates they actually have. That is the difference between a model-name swap and an accountable migration.

Run model trials with permissions, controls, and cost evidence

GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off — second editorial workflow visual

The model trial phase should not begin by asking, “Which model sounds best?” It should begin by asking, “Which replacement option is available on this product surface, under this plan, with these workspace policies, and with this exact permission profile?” OpenAI’s Help Center describes GPT-5.6 Sol as the paid-plan model for Instant and selectable thinking levels, GPT-5.6 Luna as the Free/Go default and Think model, GPT-5.6 Terra as a balanced option in Work, Codex, and the API, and GPT-6 Pro as powered by GPT-6 Astra on eligible plans. Those statements are enough to frame a trial matrix, but they are not enough to assume availability, usage limits, latency, cost, or behavior for every account.

The practical rule is simple: test only the models and controls your users can actually select in the surface being migrated. A Codex team should not approve a replacement based on a Chat-only experiment if its day-to-day risk comes from repository edits, MCP tools, terminal commands, or coding-agent review. A Work administrator should not approve a broad workspace change based only on a personal account transcript. An API team should not infer API retirement from the ChatGPT retirement notice unless OpenAI separately documents that API change; instead, it should run an API-specific evaluation against the currently documented model catalog and pricing or usage visibility available to that organization.

The official ChatGPT account states that GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex across all plans on October 14 and recommends GPT-5.6 Sol or GPT-6 Astra for Codex users. Treat that as a product-surface migration driver, not as a guarantee that any candidate model will preserve prompt behavior, tool behavior, output formatting, refusal boundaries, or cost profile. The purpose of the trial phase is to produce enough local evidence for accountable humans to choose a replacement, document accepted differences, block unsafe migrations, and roll back if production behavior diverges.

Documented comparison boundary for GPT-5.6 Sol, Terra, Luna, and GPT-6 Astra

The following table is a planning aid, not a universal model ranking. It stays inside the documented boundaries supplied by OpenAI’s Help Center and model documentation entry points, and it deliberately avoids invented performance claims. Fill the “local observation” columns during your own trial because plan eligibility, workspace policy, product surface, model picker options, and usage visibility can differ by account.

Candidate Documented product context from OpenAI sources What to verify locally before testing Migration use in the trial Do not assume
GPT-5.6 Sol OpenAI describes GPT-5.6 Sol as the paid-plan model for Instant and selectable thinking levels. Confirm whether the relevant users, workspace, or product surface can select it; confirm which reasoning or thinking controls are visible to the account. Use as a candidate for tasks where paid-plan availability and controllable reasoning levels are part of the migration path. Do not assume the same controls, limits, or behavior appear in every product, plan, region, or workspace.
GPT-5.6 Terra OpenAI describes GPT-5.6 Terra as a balanced option in Work, Codex, and the API. Confirm availability in the exact Work workspace, Codex environment, or API project being migrated. Use as a balanced trial candidate when the workload spans enterprise knowledge work, coding workflows, or API-backed automation. Do not assume “balanced” means lower risk, lower cost, faster latency, or better tool reliability without measurement.
GPT-5.6 Luna OpenAI describes GPT-5.6 Luna as the Free/Go default and Think model. Confirm whether Luna is relevant to the affected user population and whether any reasoning mode is available under their plan. Use as a candidate for user groups whose actual post-retirement default is Luna, especially where Free or Go usage is in scope. Do not assume Free/Go behavior represents Work, Codex, API, or paid-plan behavior.
GPT-6 Astra OpenAI describes GPT-6 Pro as powered by GPT-6 Astra on eligible plans. OpenAI’s Help Center also states that Astra in Codex CLI requires version 0.153.0 or later. Confirm eligible-plan access, Codex CLI version where Codex is involved, and whether Astra is selectable in the user’s surface. Use as a higher-capability candidate only where plan eligibility and operational controls are verified. Do not assume Astra is available to every user, replaces legal or professional review, or can be approved without tool and cost evidence.

Astra deserves extra care in regulated or professional settings because OpenAI’s customer stories report specific customer observations, not independent guarantees. OpenAI’s Harvey customer story says Harvey uses GPT-6 Astra to analyze, synthesize, and draft from court information, law-firm documents, case-law research, and matter context, and reports improvements in document formatting and context awareness. That is useful evidence that sophisticated legal-technology customers are evaluating Astra for document-heavy workflows, but it is not proof that your legal citations, privilege controls, client-confidentiality obligations, or jurisdiction-specific duties are satisfied. Qualified lawyer review remains mandatory before legal work product is sent, filed, relied on, or billed.

Codex teams also need to account for client version and release behavior. OpenAI’s Help Center states that Astra in Codex CLI requires version 0.153.0 or later. Separately, the stable Codex 0.158.0 release adds or changes several operational behaviors, including configurable copy-on-select and right-click paste in the fullscreen TUI, Markdown preservation when copying transcripts, support for MCP servers whose OAuth clients use pre-registered secrets through codex mcp add --oauth-client-secret, bearer-token protection for direct exec-server WebSocket connections, and default terminal input approval for elevated-permission commands. Those release notes do not make unsafe commands safe, do not automatically secure MCP clients, and do not remove the need for approval review, credential management, platform testing, and rollback.

Build a side-by-side test that measures behavior instead of impressions

A side-by-side test should compare the current GPT-5.5 baseline, where still available before retirement, with each candidate replacement that is actually reachable in the affected product surface. If GPT-5.5 is no longer available for a specific account or surface, use retained transcripts, saved outputs, test fixtures, and reviewer notes as the baseline evidence. The test must keep the input, system instructions, saved instructions, tool permissions, file set, and expected output contract as constant as the product allows; otherwise, reviewers cannot tell whether the model changed, the prompt changed, the tool set changed, or the user changed the task.

OpenAI’s evals guidance supports structured evaluation rather than anecdotal sampling. For this migration, the evaluation set should include high-frequency tasks, high-risk tasks, historically fragile tasks, safety-sensitive requests, refusal-boundary tasks, and representative tool workflows. A founder migrating customer-support automation might include refund-policy questions, ambiguous account requests, handoff triggers, and tone constraints. A developer team migrating Codex might include a small refactor, a failing-test repair, a documentation update, a dependency bump proposal, and a task that must ask for approval before running an elevated or destructive command. A legal-technology team might include document comparison, issue extraction, formatting preservation, and citation-verification prompts that require qualified legal review before acceptance.

Dimension What to measure How to capture evidence Pass condition example Blocking failure example
Quality Correctness, completeness, task fit, source use, and absence of unsupported claims. Use blind human review, rubric scoring, and retained model outputs for the same fixture. Candidate meets or exceeds baseline on critical tasks and has no new critical-error class. Candidate fabricates policy, invents citations, changes user intent, or omits mandatory constraints.
Tool calls Whether the model calls tools only when needed, uses the right tool, respects approval gates, and handles tool errors. Record tool-call logs, approved and rejected actions, arguments after redaction, and tool-error handling. Candidate uses the narrowest sufficient tool and asks for approval before consequential actions. Candidate attempts external messages, destructive commands, permission changes, purchases, submissions, or production edits without human approval.
Reasoning controls Visible reasoning or thinking settings available in the surface, and whether different settings affect answer quality, latency, or cost. Run the same fixture at each available setting and record visible settings in the evidence sheet. Selected setting produces acceptable results within latency and usage expectations. Team approves a setting that is unavailable to the target users or materially too slow for the workflow.
Latency Time to first useful response, time to final answer, and delay added by tool calls or approvals. Use application logs, user-observed timings, or product-visible timestamps where available. Latency remains within the workflow’s agreed threshold for the user group. Critical workflows time out, miss service targets, or cause users to bypass review.
Usage Message consumption, quota usage, API token usage, tool invocations, or other usage indicators visible to the organization. Capture plan-visible, workspace-visible, or API-visible usage metrics without inventing missing values. Usage remains within team allowance, budget, or capacity planning assumptions. Candidate repeatedly consumes disproportionate usage for routine tasks or exhausts allowances during normal volume.
Cost where visible API cost, workspace cost signals, or internal allocation metrics that the account can legitimately observe. Use official billing views, API usage dashboards, project-level logs, or approved FinOps exports. Projected cost is approved by the budget owner with documented assumptions. Cost is unknown for a high-volume workflow and the team wants to migrate without a cap or monitoring plan.
Refusals Whether the model refuses prohibited requests, complies with allowed requests, and gives safe alternatives where appropriate. Include benign, ambiguous, and disallowed fixtures; classify refusal accuracy, over-refusal, and under-refusal. Candidate refuses disallowed requests and does not block ordinary allowed business tasks without reason. Candidate provides harmful instructions, bypass guidance, or unsafe operational detail; or it refuses required benign tasks at unacceptable rates.
Safety Privacy preservation, secret handling, youth-safety behavior, regulated-domain caution, and escalation to qualified humans. Use redacted fixtures and policy rubrics; never include real credentials, account numbers, private health data, or unnecessary personal data. Candidate asks for minimal information, avoids exposing secrets, and flags consequential decisions for human review. Candidate asks users to paste secrets, private identifiers, privileged files, or sensitive data not needed for the task.
Output compatibility Schema conformance, formatting, Markdown, JSON validity, field names, tone, citations, and downstream parser compatibility. Run outputs through validators, snapshot tests, parser tests, and human review for narrative formats. Candidate output passes downstream checks or differences are documented and accepted. Candidate breaks automation by changing keys, nesting, citation style, headings, or required sections without approval.

The table should be converted into a test sheet for each workflow owner. Do not aggregate away critical failures by averaging. A candidate that scores well on marketing drafts but sends an external customer message without approval in one fixture has failed that workflow until the tool permissions, prompt contract, or approval gate is corrected and retested.

Use an explicit trial matrix for product, plan, and workspace reality

A common migration failure is testing a model in the wrong place. The same model name can appear across products while surrounding controls differ, and OpenAI’s Help Center explicitly distinguishes Chat, Work, Codex, and API contexts. Work and Codex allowances are separate from Chat, so a Chat transcript alone does not prove Codex readiness, and a Codex success does not establish API throughput, billing, or integration behavior. Your test matrix should therefore list the exact product surface, account class, workspace, project, client version, and policy setting for every result.

Recommended trial matrix fields:

workflow_id: "codex-refactor-small-service"
owner: "engineering-platform"
surface_under_test: "Codex CLI"
codex_cli_version: "record actual version from /status or approved inventory"
candidate_model: "GPT-5.6 Terra"
availability_verified_by: "named workspace or platform admin"
reasoning_control_visible: "record exact visible option or 'not visible'"
tool_profile: "read-only repo tools; tests require approval; no production deploy tool"
connectors_enabled: "list approved connectors only"
fixture_id: "repo-fixture-014"
baseline_evidence: "gpt-5.5 transcript or retained output path"
latency_observation: "record measured local value"
usage_observation: "record visible usage metric or 'not visible in this surface'"
cost_observation: "record official billing metric or 'not visible to evaluator'"
reviewer_decision: "pass / pass with conditions / fail"
blocking_issue_id: "if any"
rollback_required: "yes / no"
signoff_owner: "named accountable owner"

This structure prevents a senior sponsor from approving a migration based on a demo that did not include the actual connectors, actual permissions, actual client version, or actual user plan. It also helps administrators answer a later incident question: “Was this behavior tested in the environment where it occurred?” If the answer is no, the migration evidence package was incomplete.

Control reasoning settings as a change variable

Where a product surface exposes selectable thinking or reasoning levels, treat the setting as part of the configuration, not as a casual user preference. OpenAI describes GPT-5.6 Sol as supporting Instant and selectable thinking levels for paid plans, but the Help Center statement does not mean every workspace exposes the same controls or that every workflow should use the deepest available setting. Higher-effort reasoning may be useful for complex synthesis, code planning, legal-document comparison, or multi-step analysis, while a fast setting may be more appropriate for classification, routing, or short drafts. The trial must measure the setting actually intended for production use.

The safest procedure is to run each representative fixture under the visible settings that the target users can select, then record the tradeoff in the evidence sheet. If the quality gain appears only at a setting that causes unacceptable latency, usage consumption, or reviewer fatigue, the team should either narrow that setting to specific workflows or choose a different model. If a model produces acceptable quality only when a hidden or unavailable control is used by the migration team, the result is not valid for the affected users.

Operational rule: a model plus its reasoning setting plus its tool profile equals the tested configuration. Changing any one of those three after sign-off requires either retesting or an explicit risk acceptance by the accountable owner.

For API evaluations, use OpenAI’s model documentation and current API guidance as the factual boundary, and do not copy assumptions from ChatGPT or Codex. Record request parameters, output contracts, validators, and usage telemetry approved for the project. If the migration team cannot see cost, latency, or usage in a trustworthy way, the correct conclusion is “not yet measurable,” not “no impact.”

Require least-privilege tool permissions before judging model quality

Tool permissions can make a candidate look better or worse for reasons unrelated to model quality. A model with broad connectors may answer a knowledge question better because it can retrieve more context, while the same configuration may create an unacceptable privacy or compliance risk. A coding model with shell access may fix a test faster, but if it can run elevated commands without appropriate approval, the speed improvement is not a migration success. Review permissions before and during the trial, and retest after any permission change.

Least privilege means the model gets only the tools, connectors, repositories, files, domains, scopes, and command classes required for the fixture. It does not mean “give broad access for the test and tighten later,” because broad-access tests produce evidence for a configuration you should not deploy. For Codex and MCP workflows, review server-level approval modes, per-tool approval modes, enabled and disabled tool lists, OAuth or bearer-token authentication, sandbox roots, writable paths, network boundaries, and whether project-scoped configuration is loaded only for trusted projects. OpenAI’s Codex configuration reference states that project configuration cannot override machine-local provider, auth, or telemetry keys, which is an important boundary but not a substitute for policy review.

Permission area Minimum migration review Evidence to retain Stop condition
Connectors and knowledge sources Confirm each connector is necessary for the fixture, authorized for the user group, and limited to appropriate data. Connector list, owner approval, data-classification note, and screenshots or exports where policy allows. Connector exposes confidential, regulated, privileged, or youth-related data not needed for the tested task.
MCP servers and tools Review server identity, OAuth or bearer-token configuration, callback registration, enabled tools, disabled tools, and approval mode. Redacted MCP configuration, approval policy, callback verification, and tool-call logs. Unknown server, broad tool scope, missing approval, unverified callback, or any request to paste secrets into prompts.
Repository access Limit to repositories needed for the fixture; use Git checkpoints before and after tasks as OpenAI’s Codex CLI guidance recommends. Repository list, branch name, checkpoint references, diff summaries, and reviewer notes. Agent attempts to alter unrelated repositories, generated files outside scope, or protected branches without approval.
Terminal commands Classify read-only, build/test, network, elevated, destructive, and production-impacting commands; require approval for elevated and consequential actions. Command transcript, approval decisions, sandbox profile, and redacted logs. Command would delete data, change permissions, deploy, exfiltrate data, or alter production without explicit human approval.
External communication Disable or require approval for sending emails, tickets, messages, posts, filings, submissions, or customer-visible content. Tool policy, approval record, draft review, and final human approver. Model attempts to send or publish externally without human review and approval.
Secrets and credentials Use approved secret storage and placeholders; never paste tokens, keys, passwords, OAuth client secrets, or private keys into prompts or reports. Statement of secret-handling method, redaction proof, and owner approval without revealing the secret. Any test step requires exposing real credentials to the model, transcript, repository, ticket, or migration report.

OpenAI’s Codex 0.158 release notes add support for MCP servers whose OAuth clients use pre-registered secrets through codex mcp add --oauth-client-secret. In a migration trial, that capability should be handled by platform owners using approved secret handling, not by asking developers to paste a real client secret into a model prompt or shared document. The safe evidence is that the secret was configured through an approved process and that the tool worked under the narrow intended scope; the unsafe evidence is a transcript, ticket, shell history, or pull request that contains the secret itself.

Connector review must include data classification and downstream action risk

Connector review is not just an access checklist. It should combine data classification, user authorization, intended purpose, output exposure, retention expectations, and downstream action risk. A connector to an internal knowledge base may be acceptable for policy summarization but unacceptable for drafting individualized HR decisions without human review. A code-hosting connector may be acceptable for reading a test fixture but unacceptable for changing protected branches. A legal research or matter-document connector may be acceptable for issue spotting under attorney supervision but unacceptable for autonomous legal conclusions or external filings.

Use a two-step connector approval. First, run read-only discovery with redacted or non-sensitive fixtures to confirm that the candidate can use the connector appropriately. Second, run a controlled representative fixture with the minimum real context necessary and a named reviewer. Do not include unnecessary private data, credentials, account numbers, health information, protected student information, privileged documents, or confidential client materials in a test unless the organization’s policy, contract, and professional obligations allow that use and a qualified owner approves it.

Recommended connector review questions:

1. What exact workflow requires this connector?
2. Which user group is authorized to use it?
3. What data classes can the connector expose?
4. Can the same evaluation be run with redacted or synthetic data first?
5. Which actions are read-only, draft-only, or externally consequential?
6. Which tool calls require human approval before execution?
7. What evidence will be retained without exposing secrets or private content?
8. Who can disable the connector if the trial fails?
9. What is the rollback path if the connector changes output behavior?
10. Which policy owner signs off before broader rollout?

For youth-facing, education, health, finance, legal, and employment workflows, treat connector access as a compliance decision rather than a convenience feature. The model should not be used to infer eligibility, discipline, diagnosis, creditworthiness, legal rights, safety status, or other consequential outcomes without qualified human governance, appropriate policy controls, and a documented review path. A migration deadline does not reduce those obligations.

Measure cost, usage, and latency without inventing missing data

Cost analysis should use only information visible through official billing, usage, product, or administrative surfaces available to the organization. Some ChatGPT or Work trials may not expose granular per-task cost to the evaluator. Some API projects may expose usage more directly. Codex and Chat allowances are separate according to OpenAI’s Help Center, so a migration team should not treat a low Chat usage pattern as proof that Codex capacity is sufficient. If cost or usage is not visible, record it as unavailable and require a budget owner or administrator to provide an approved measurement method before high-volume rollout.

For API-backed workflows, capture input size, output size, tool calls, retries, validation failures, and any project-level usage or billing data the organization is authorized to view. For Work and Codex workflows, capture visible allowance consumption, number of interactions per task, number of tool calls, number of approvals, time spent by reviewers, and rework rate. Human review time is a real cost even when the product bill is unchanged, and a model that produces more review churn may be more expensive operationally than a model with a higher apparent capability but a lower correction burden.

Cost or capacity factor Why it matters How to measure conservatively Approval owner
Visible product or API usage Shows whether the migration fits within known allowances, limits, or billing expectations. Use official dashboards, project reports, or workspace admin views; do not estimate from memory. Workspace admin, API project owner, or FinOps owner.
Retries and failed validations Repeated attempts can increase usage, latency, and reviewer fatigue. Count retries per fixture and classify the cause: model error, prompt ambiguity, tool failure, or validator strictness. Workflow owner and evaluation lead.
Tool-call volume Tools may add latency, permission risk, audit requirements, and third-party service usage. Log each tool call with redacted arguments, approval status, outcome, and error handling. Security owner, platform owner, or connector owner.
Reviewer time Human review is mandatory for consequential actions and can dominate migration cost. Record minutes spent reviewing each output, diff, draft, or action proposal. Team manager or professional reviewer.
Latency and queueing Slow responses can break support, coding, education, operations, or customer workflows. Measure start time, first useful response, final answer, and approval delays for each fixture. Product owner or service owner.
Rollback and remediation effort A risky migration may create hidden support, incident, and rework costs. Estimate only from documented trial failures, required prompt changes, permission changes, and retesting hours. Migration sponsor and risk owner.

Do not normalize cost by choosing only short, easy prompts. Include long-context prompts, file-heavy tasks, tool-using tasks, safety-boundary tasks, and output-validation tasks that reflect actual production work. A model that is inexpensive on short drafts may be costly when it repeatedly breaks JSON, overuses connectors, or requires multiple human review cycles for the same consequential output.

Test output compatibility before users discover broken workflows

Output compatibility is often the first production failure after a model migration. A replacement model may be more fluent but still break a parser, change a heading hierarchy, rename a JSON field, omit a required disclaimer, alter a citation format, or produce a longer explanation than a downstream system can accept. For developers and administrators, this means schema validation must be part of the side-by-side trial. For knowledge workers and educators, it means the rubric, tone, and formatting expectations must be documented before review. For legal-technology teams, it means citation format, issue order, source references, and lawyer-review markers must be checked before any draft is treated as useful work product.

A compatibility test should include both machine validators and human reviewers. Machine validators can check JSON validity, XML structure, required fields, maximum length, prohibited fields, citation placeholders, or Markdown structure. Human reviewers should check whether the answer still satisfies the business purpose, avoids unsupported claims, and preserves required cautions. If the model output must feed a workflow engine, ticketing system, gradebook, document assembly system, code generator, or legal drafting pipeline, run the candidate output through that downstream system in a controlled environment before sign-off.

Sample output-compatibility gate:

A candidate model cannot pass this workflow unless:
- every required field is present;
- no undocumented field is introduced;
- values match the expected type and allowed range;
- citations or source references use the required format;
- the answer includes required

Rollout design: shadow first, canary second, broad migration last

A safe GPT-5.5 transition should move through three operational states: shadow runs, canary use, and broad rollout. A shadow run is a non-production comparison in which the team sends the same representative task fixture to the current workflow and to the candidate replacement workflow, then records differences without letting the new output affect customers, codebases, documents, tickets, legal positions, grades, financial decisions, or external messages. The point is not to prove behavioral equivalence; the point is to discover the failure modes that matter before users depend on the new model.

Shadow runs are especially important because OpenAI’s notice concerns GPT-5.5 retirement across ChatGPT, ChatGPT Work, and Codex product surfaces, while the official documentation for available models, Work, Codex, and ChatGPT behavior can vary by plan, product, workspace permissions, region, and current documentation. Treat every product surface as a separate migration lane. A prompt that works acceptably in a personal ChatGPT session may behave differently inside a workspace with different instructions, connectors, tool permissions, data controls, or Codex configuration.

A practical shadow run uses frozen fixtures rather than ad hoc user anecdotes. For example, a software team can replay ten representative bug-fix tasks, five dependency-update tasks, five code-review tasks, and five documentation tasks against the candidate Codex model while preserving Git diffs, tool calls, approval requests, test results, and reviewer notes. A knowledge-work team can replay policy summaries, spreadsheet explanations, meeting-note transformations, research outlines, and customer-response drafts, then score factuality, omission risk, tone, output structure, and escalation behavior.

Run canaries only after the shadow data shows that the replacement path is usable under explicit acceptance thresholds. A canary user is not merely an enthusiastic early adopter; a canary is a trained participant operating inside a bounded scope, with known workflows, escalation instructions, rollback access, and an obligation to report defects. Choose canary users who can recognize bad outputs, understand the business process, and stop a workflow when the model creates risk. Do not choose canaries solely because they are heavy users; high volume without disciplined reporting can hide the signal you need.

For Codex canaries, the canary boundary should include repository scope, branch rules, tool permissions, network assumptions, required Git checkpoints, and reviewer approval for code changes. OpenAI’s Codex and Work documentation should be treated as the current source for product behavior, and teams should confirm model availability and permissions in the actual environment being migrated. Do not assume that a recommendation for Codex users to consider GPT-5.6 Sol or GPT-6 Astra means the same model, reasoning controls, or allowances are available in every account or workspace.

Rollout state Who participates What is allowed Required evidence Exit rule
Shadow run Migration team, reviewers, application owners Offline comparison, fixture replay, non-production Codex trials, redacted document tasks Side-by-side outputs, score sheets, error taxonomy, cost and latency notes, tool-permission observations Critical errors are understood, mitigations are documented, and canary scope is approved
Canary Named trained users in a bounded workflow Limited real work with human review before external or consequential action User reports, sampled transcripts or redacted artifacts, approval logs, defects, rollback tests Launch gates are met for a defined duration or task count without unresolved stop conditions
Broad rollout Approved user groups or workspace segments Normal use within policy, permissions, and review requirements Communication records, sign-off, monitoring plan, support queue, post-migration sampling schedule Migration owner confirms GPT-5.5-dependent workflows have replacement paths or documented exceptions

Communication plan for users, reviewers, and support teams

Migration communication must tell users what is changing, what is not known, what they must do differently, and where to report failures. Avoid vague announcements such as “we upgraded the model.” State the affected product surface, the retirement date from OpenAI’s ChatGPT notice, the replacement path being tested, and the local rules for sensitive data, external messages, code changes, and approvals. The communication should also state that output may change and that the organization is not promising exact behavioral equivalence.

Segment communication by role. Developers need guidance on Codex model selection, repository checkpoints, review expectations, and tool permissions. Enterprise administrators need workspace readiness, support escalation, usage observation, and exception handling. Legal-technology professionals need source-verification requirements, confidentiality reminders, citation review, and qualified-lawyer approval. Educators need policies for student work, grading assistance, age-appropriate use, and academic-integrity review. Parents and guardians need conservative youth-safety boundaries and a reminder that model responses are not a substitute for qualified support in serious situations.

A support-ready announcement should include a defect-report template. Require users to provide the task category, product surface, model shown if visible, workspace or project context, non-sensitive prompt summary, expected output contract, actual failure, severity, and whether the task involved tools, connectors, files, code, or external communication. Do not ask users to paste secrets, personal identifiers, confidential legal facts, private health information, payment details, privileged material, or restricted customer data into support forms unless an approved secure process exists for that material.

Recommended communication template:

Subject: GPT-5.5 migration readiness for [team/workspace/product surface]

OpenAI has stated that GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex across all plans on October 14. Our team is migrating affected workflows through shadow testing, canary use, and sign-off before broad rollout.

What changes for you:
- Use the approved replacement path for [workflow/product surface].
- Do not assume outputs will match GPT-5.5.
- Keep prompts, output schemas, and workflow instructions in the approved location.
- Get human approval before external messages, code merges, submissions, purchases, bookings, permission changes, publication, legal commitments, or other consequential actions.

How to report an issue:
- Provide the task category, product surface, non-sensitive prompt summary, expected output, actual failure, and severity.
- Do not include credentials, tokens, private keys, personal identifiers, confidential client facts, health information, payment details, or privileged material in the report.

Rollback:
- If a stop condition occurs, pause the workflow and notify [migration owner] and [support channel].

The support desk should receive a different message than end users. Give support teams a triage map that distinguishes product availability questions, model-behavior defects, permission failures, connector or tool issues, cost anomalies, latency incidents, and policy escalations. Support staff should know when to send a problem to workspace administration, security, legal, engineering, finance, or the migration owner. Without this routing, canary users will report issues, but the organization will not convert those reports into launch decisions.

Launch gates and stop conditions

A launch gate is an explicit condition that must be true before the migration advances. Launch gates should be written before canary testing begins, because teams under retirement pressure may otherwise reinterpret mixed evidence as “good enough.” Each gate should name an owner, the evidence required, the decision authority, and the product surface to which it applies. A gate for ChatGPT Work document drafting should not automatically approve Codex repository changes, and a gate for one region or workspace should not automatically approve another.

The first launch gate is inventory closure. Every known GPT-5.5 dependency must be classified as migrated, not affected, retired, or exception-managed. Exception-managed means the owner has documented why a workflow cannot be fully migrated before the date, what temporary control will be used, who accepts the risk, and how users will be prevented from unknowingly relying on the retired path. “Unknown” is not an acceptable production state for a high-value workflow.

The second launch gate is representative evaluation completion. OpenAI’s evals guidance supports the broader principle that model changes should be tested against representative tasks rather than judged by isolated examples. Your local gate should require enough fixture coverage to represent high-frequency tasks, high-risk tasks, edge cases, and known failure modes. The gate should include pass/fail thresholds, not just reviewer comments. For example, a legal drafting workflow might require no unsupported citations in the reviewed sample, correct preservation of defined terms, and mandatory escalation when source material is insufficient.

The third launch gate is permission review. Tool permissions, connectors, file access, repository access, browsing behavior, execution permissions, and MCP or Codex configuration should be reviewed before broad release. A model migration can change how often a tool is requested, how instructions are interpreted, or how users rely on automation. The safe default is least privilege, with human approval for elevated, external, destructive, financial, legal, administrative, or publication actions.

The fourth launch gate is cost and latency visibility. Do not invent price, limit, or throughput assumptions. Instead, record observed usage and latency during shadow and canary runs in the actual product surface and plan where the workflow will operate. If official documentation changes, re-check the assumptions. A workflow can be functionally acceptable but operationally unsuitable if response time breaks a support process, if usage patterns exceed a workspace policy, or if the cost owner cannot forecast the change.

The fifth launch gate is rollback proof. The rollback plan must be executable by named owners, not merely documented in a planning slide. If GPT-5.5 is no longer available in a product surface after the retirement date, rollback cannot mean “switch back to GPT-5.5” unless OpenAI’s current product behavior explicitly permits that option. Rollback may instead mean reverting prompts, disabling a tool path, routing work to a different approved model, pausing a connector, restoring a previous automation version, returning to manual review, or narrowing the user population.

Stop condition Example signal Immediate action Decision owner
Critical safety or compliance failure Model drafts prohibited advice, exposes restricted content, or bypasses required escalation Pause affected workflow, preserve evidence, notify security/compliance owner Risk owner and migration owner
Unauthorized tool or data access attempt Workflow requests a connector, repository, command, or file class outside approved scope Block or revoke the tool path, review permissions, rerun fixture Workspace administrator or platform owner
Output contract break JSON, citation format, document structure, or code patch format becomes incompatible Hold rollout for that workflow, repair prompt/schema, retest Application owner
Unacceptable factuality or reasoning defect Reviewed sample exceeds the pre-set threshold for omissions, hallucinations, or unsupported claims Return to shadow testing and revise task routing or instructions Workflow owner and subject-matter reviewer
Cost, usage, or latency anomaly Observed operation no longer fits the team’s budget, timing, or workspace policy Pause expansion, analyze logs, adjust routing or scope Finance owner and product owner

Rollback ownership when the old model is retiring

Rollback planning is harder in a retirement migration because the previous state may disappear from the product surface. That is why rollback ownership must be assigned before October 14, not discovered during an incident. The owner should have authority to pause the rollout, disable an affected automation, change workspace instructions, restrict a connector, revert a prompt package, or route users to a manual process. If the owner can only “file a ticket,” the rollback plan is incomplete.

Separate rollback ownership into four layers. The workflow owner decides whether the business process can continue. The platform or workspace administrator changes permissions, model settings, connectors, or Codex configuration where available. The security or compliance owner determines whether evidence preservation, incident response, or user notification is required. The executive or accountable service owner accepts residual risk when the migration proceeds under an exception.

A useful rollback plan contains a decision tree rather than a single button. If a prompt regression breaks a formatting contract, revert the prompt package and rerun the fixture. If a connector permission is too broad, disable the connector or reduce the tool allowlist and rerun the affected tasks. If Codex proposes unsafe commands, pause the repository canary, require manual implementation, and review approval policy. If a high-risk document workflow produces unsupported assertions, route the work to subject-matter review or suspend the AI-assisted path until the evaluation is repaired.

Do not instruct users to preserve continuity by copying sensitive data into unapproved accounts, personal tools, or unmanaged prompts. Retirement pressure does not justify bypassing workspace permissions, data-handling policy, legal privilege controls, school policy, parental controls, contractual obligations, or security review. If a critical workflow cannot be migrated safely by the retirement date, the conservative option is a documented manual fallback with accountable approval.

Rollback record template:

Workflow:
Product surface:
Replacement model/path:
Rollback owner:
Backup owner:
Stop condition triggered:
Evidence location:
Immediate containment action:
User communication required: yes/no
Permission change required: yes/no
Manual fallback:
Retest fixture:
Approval required to resume:
Final disposition:

Post-migration sampling and drift detection

Migration does not end on launch day. Post-migration sampling checks whether real usage still resembles the fixtures and canary evidence. Sampling should begin immediately after broad rollout and continue through at least one normal business cycle for the workflow. For a support team, that may mean sampling tickets across weekdays and weekend shifts. For finance or legal operations, it may mean sampling month-end, quarter-end, or matter-specific work. For education, it may mean sampling assignment drafting, rubric assistance, and student-support scenarios separately.

Use risk-weighted sampling instead of uniform sampling. High-consequence workflows deserve more review than low-risk brainstorming. Sample external messages before they are sent, code changes before they are merged, legal or policy drafts before they are relied on, and administrative actions before permissions or records are changed. For low-risk internal ideation, lighter sampling may be acceptable if users know the boundaries and escalation rules.

Post-migration sampling should record both model-output defects and process defects. A model-output defect includes an unsupported claim, omitted constraint, wrong format, poor tool choice, or unsafe recommendation. A process defect includes missing human approval, users selecting the wrong model or workspace, failure to preserve evidence, untracked prompt edits, unreviewed connector access, or support tickets without enough detail to reproduce the problem. In many migrations, process defects create more risk than the model change itself.

Maintain a drift register for changes after sign-off. Availability, model choices, reasoning controls, workspace permissions, and documentation can change over time, and OpenAI’s official model and product pages remain the source to re-check. Add entries when a prompt changes, a connector is added, a repository permission changes, a workspace policy changes, a new model option becomes available, a model option disappears, or a user group expands. Each entry should name the owner, the expected impact, and whether representative fixtures must be rerun.

Sampling item What to inspect What to record Escalation trigger
Prompt adherence Whether the output follows system, workspace, and workflow instructions Prompt version, output artifact, reviewer score, defect class Repeated instruction drift or a critical instruction violation
Tool use Whether requested tools match approved permissions and task intent Tool name, purpose, approval status, result, blocked attempts Unexpected tool request or overbroad access pattern
Output contract Whether format, schema, tone, citation structure, or code patch format remains compatible Contract version, pass/fail result, parser or reviewer notes Broken downstream automation or human review bottleneck
Human approval Whether consequential actions were reviewed before execution Approver, timestamp, scope, decision, exception notes Action taken without required approval
Cost and latency Whether observed operation still fits budget and service expectations Observed usage, timing, product surface, relevant plan context Budget owner concern, support delay, or workspace-limit friction

Retirement-day readiness checklist

Retirement-day readiness is a practical operations exercise, not a ceremonial announcement. The migration owner should verify that every affected team knows what to use instead of GPT-5.5, where to report issues, and what to do if a replacement path fails. The support team should have the triage map, the administrator should know the permission-change process, and workflow owners should have authority to pause risky tasks without waiting for a committee meeting.

On the day before retirement, freeze non-essential prompt changes for high-risk workflows. This does not mean no emergency fix can be made; it means teams should avoid introducing unrelated prompt rewrites, connector additions, schema changes, or permission expansions that make retirement-day failures hard to diagnose. If a change is necessary, record the reason, owner, evidence, and rollback path.

On retirement day, verify the actual user experience in each product surface. Ask named testers to confirm the approved path in ChatGPT, ChatGPT Work, and Codex as applicable to the organization. If a model selector, default model, reasoning control, or Codex option differs from the migration plan, record what is visible and compare it against current official documentation and workspace policy. Do not assume a discrepancy is an OpenAI error or a local error until the environment, plan, permissions, rollout status, and documentation have been checked.

Run a small smoke test for each critical workflow. A smoke test is not a full re-evaluation; it is a minimal confirmation that the approved replacement path still works. For Codex, that may include opening a trusted repository, confirming the intended model path if visible, running a read-only analysis task, observing tool approval behavior, and confirming no unexpected permission request appears. For a document workflow, it may include a redacted fixture, required output structure, citation or source-handling rule, and reviewer sign-off.

  1. Confirm affected product surfaces: ChatGPT, ChatGPT Work, Codex, or a documented non-affected lane.
  2. Confirm each workflow has an approved replacement path or an exception-managed manual fallback.
  3. Confirm prompt packages, instructions, output schemas, and fixture results are versioned.
  4. Confirm tool permissions, connectors, repositories, files, and execution paths remain least-privilege.
  5. Confirm users received communication and support knows the triage process.
  6. Confirm stop conditions and rollback owners are visible to canary leads and administrators.
  7. Confirm cost and latency observations are available to the product or finance owner.
  8. Confirm high-risk external or consequential actions still require human approval.
  9. Confirm post-migration sampling starts immediately after broad rollout.

Accountable sign-off package

Accountable sign-off should be a written acceptance of evidence, scope, residual risk, and operating controls. It is not a generic approval to “use the new model.” A useful sign-off package contains the migration charter, dependency inventory, representative evaluation results, side-by-side review notes, prompt and schema versions, tool-permission review, cost and latency observations, canary evidence, support plan, rollback plan, exception list, and post-migration sampling schedule.

Require at least three kinds of sign-off for material workflows. The workflow owner confirms the replacement path supports the business task. The platform or workspace owner confirms configuration, permissions, and supportability. The risk owner confirms that privacy, security, legal, safety, or compliance controls are acceptable for the defined scope. For software delivery, add engineering approval for repository and deployment processes. For education, add academic-policy approval where AI assistance affects assignments, grading, or student-facing guidance. For legal-technology use, require qualified lawyer review for legal conclusions, filings, client advice, citations, and matter-specific strategy.

The sign-off should also identify what is not approved. For example, a team may approve internal summarization but not external customer replies; code explanation but not autonomous code merge; first-draft policy assistance but not final compliance interpretation; legal research organization but not filing-ready legal analysis; lesson planning but not unsupervised student counseling. Clear exclusions prevent users from expanding a migration approval into unreviewed use cases.

Accountable sign-off statement:

I approve the GPT-5.5 migration for [workflow] on [product surface] using [approved replacement path], limited to [scope]. I have reviewed the dependency inventory, representative evaluation results, prompt and output-contract versions, tool-permission review, cost and latency observations, canary evidence, stop conditions, rollback plan, and post-migration sampling schedule.

This approval does not authorize [excluded uses]. Human approval remains required for external messages, submissions, payments, purchases, bookings, destructive actions, permission changes, publication, legal commitments, production deployment, and other consequential operations.

Workflow owner:
Platform/workspace owner:
Security/compliance/legal owner:
Date:
Review date:

Do not let sign-off depend on undocumented personal confidence in a model. The artifact should be clear enough that a new administrator, auditor, support lead, or successor owner can understand why the migration was approved and what evidence would require reopening the decision. If official OpenAI documentation changes or a product surface behaves differently from the signed plan, treat that as a change event and decide whether fixtures, permissions, or user communications must be updated.

Conclusion: migrate the operating system around the model, not just the model name

The GPT-5.5 retirement notice creates a calendar deadline, but the migration risk comes from hidden dependencies: prompts that nobody versioned, output contracts that downstream tools silently expect, connectors that expose more data than the task needs, Codex permissions that differ by repository, and users who assume a new model will behave like the old one. OpenAI’s official materials identify the affected ChatGPT, ChatGPT Work, and Codex context and point users toward current model options, but they do not remove the need for local evaluation, permission review, cost observation, rollback planning, and accountable acceptance.

The safest organizations will treat the transition as a controlled change program. They will run shadow comparisons before canaries, select trained canary users before broad rollout, define launch gates before interpreting evidence, assign rollback owners before the old path disappears, sample real work after migration, and require human approval for consequential actions. They will also keep checking current OpenAI documentation because availability, model choices, reasoning controls, and usage conditions can depend on plan, product, workspace permissions, region, and ongoing product changes.

The final decision rule is simple: do not ask whether the replacement model is “better” in the abstract. Ask whether the approved replacement path, in the actual product surface, with the actual permissions, prompts, users, data boundaries, output contracts, costs, and human review rules, is acceptable for the defined workflow. If the answer is documented, tested, and signed by accountable owners, the migration is operationally ready. If the answer depends on hope, memory, or informal impressions, keep it in shadow or canary until the evidence is strong enough to carry the workflow.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library →

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this