25 Codex Prompts for Open-Source Package Maintainers: From Issue Intake to a Safe, Evidence-Backed Release

A maintainer guiding a small package-shaped object through five illuminated gates toward a sealed release crate
A maintainer guiding a small package-shaped object through five illuminated gates toward a sealed release crate
The playbook treats maintenance as a gated evidence journey in which authority, plan, patch, release class, and publication each require human control.

Evidence checkpoints

Documented point: Current page accessed 2026-10-02 says Codex command-line interface (CLI)A text-based interface for running commands and tools. Open glossary entry can inspect, edit, and run local code; support repeatable automation; and perform dedicated reviews without changing the working tree. Local tools and permissions are user-controlled; this is not release authorisation or a guarantee of correct code. [official source 1]

Documented point: Current page accessed 2026-10-02 says published environments reuse prepared filesystem state and each new task has its own workspace; configured network secrets are substituted for allowed Hypertext Transfer Protocol Secure (HTTPS)The encrypted form of web communication protected with Transport Layer Security. Open glossary entry destinations. Repository access, network policy, credentials, workspace controls, and service permissions remain separate; never solicit raw secrets. [official source 1]

Documented point: OpenAI support article updated 2026-10-02 distinguishes Chat, Work, and Codex and says Codex is for software development/technical work; availability depends on plan, workspace settings, and rollout. Do not conflate Codex with normal ChatGPT Chat/Work, local desktop execution, Cloud, or application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry access. [official source 1]

Documented point: Current guide says effective Codex prompts include goal, context, constraints, and done-when, and recommends planning, testing, and reviewing changes. The article’s fifth fixed output-artifact component is an editorial device, not a required OpenAI prompt format. [official source 1]

Documented point: Current guide says Codex reads AGENTS.md before work, layers instruction files from repository root toward the current directory, and has a default 32 KiB combined-guidance limit. Instructions can be stale, conflicting, incomplete, or truncated; they do not guarantee safe or correct output. [official source 1]

Documented point: Current guidance separates sandbox technical boundaries from approval policy and describes workspace-write plus on-request as a lower-risk local automation preset. Never prescribe danger-full-access as a default; platform and managed-workspace rules vary. [official source 1]

Maintaining a public dependency creates an asymmetric responsibility. A reporter can describe one surprising result and walk away; the maintainer must decide whether it is reproducible, whether it violates a documented contract, which supported consumers could be affected, and whether any change would create a second breakage. Codex can help inspect evidence and prepare reviewable artefacts, but it does not inherit the project’s authority. A generated patch, passing test or favourable review remains evidence for a maintainer—not permission to merge or release.

This playbook follows a fictional issue in Northstar SDK, an imaginary open-source software development kit. Version 2.5.0 is reported to reject the compact Coordinated Universal Time (UTC)The internationally agreed standard for world time, used here to identify the time reference in timestamps and offsets. Open glossary entry-offset input parseTimestamp("2026-10-02T09:30:00+0000"), while version 2.4.2 is reported to accept the same synthetic timestamp. The reporter expects the public parser to accept this compact +0000 form as it accepts the colon-separated +00:00 form. These names, versions, inputs and reported results are illustrative only, not observations about a real package, test or report.

The evidence chain is deliberately stricter than “ask for a fix”. It records the supplied report, applicable repository instructions, project authority, an approved reproduction design, observed command results and the public contract. Later stages can use those records to assess compatibility, propose a patch and prepare a release hand-off. At every stage, confirmed evidence is separated from assumptions, generated interpretations and decisions reserved for authorised humans.

What this workflow automates—and what remains evidence work

Work only inside an authorised repository with approved sandbox and network permissions. Keep a release evidence ledger of source links, reviewed diffs, tests, unresolved findings and the named human maintainer’s sign-off; these proposed prompts are not proof that a package was tested or released.

Codex CLI can inspect, edit and run code in a local repository. OpenAI’s Codex CLI documentation, accessed 2 October 2026, also describes repeatable automation and a dedicated review mode that does not change the working tree. Those capabilities support repository discovery, command execution and structured review. They do not establish that a report is true, a contract has been broken or a release is safe.

The practical unit of automation should therefore be an artefact with provenance. For issue intake, that might be a ledger linking each assertion to the issue text or a repository path. For reproduction, it should be a command record containing the exact fixture, environment and observed output. A maintainer can reject, amend or approve each artefact without surrendering control of the repository.

Use the sequence selectively. A documentation question may stop after classification and contract tracing. A reproducible regression can continue into compatibility analysis and a scoped patch. A possible private vulnerability must leave the public workflow immediately and follow the project’s authorised disclosure process. Do not paste private vulnerability details, credentials, customer logs or personal information into a prompt.

A useful operating rule is: delegate collection and drafting when the result can be independently inspected; retain decisions where the result changes public obligations. Repository search and test-log formatting are suitable delegated tasks. Deciding support policy, semantic versioning (numbering major, minor and patch releases according to compatibility), deprecation, disclosure, merge timing or publication is not.

What Codex cannot authorise

Codex cannot grant repository permission, reinterpret a project’s governance rules, accept contributor terms, waive branch protection or approve a release. It also cannot decide that an external service may be contacted merely because a network route exists. Access to a filesystem, repository, credential, network destination and service operation are separate grants.

The same separation applies to consequential decisions. Human review is required for security, privacy, money, employment, government and other high-impact matters. In package maintenance, this includes vulnerability handling, paid-service access, contributor sanctions, licence disputes, disclosure timing and changes that could affect regulated or safety-critical users. Route such matters to the authorised people and policies rather than asking Codex to settle them.

OpenAI’s GitHub integration guidance says repository review instructions do not replace tests, branch protections or required approvals. Apply that principle throughout this playbook: generated findings can inform a reviewer, but cannot self-approve. If a task reaches merge, tag, registry publication, release signing or credential use, stop for an authorised maintainer.

Boundary What Codex may do within an approved task What must be established separately Decision rule
Repository authority Read authorised files; propose edits on an isolated branch or worktree; inspect a supplied issue or pull request. Ownership, contributor permissions, protected-branch rules and permission to modify the repository. If authority is absent or ambiguous, restrict work to read-only analysis and ask a maintainer.
Network access Use destinations explicitly allowed by the configured environment and task. Permission to contact each service, send particular data and perform the requested operation. A reachable endpoint is not an authorised endpoint. Stop before any unapproved request.
Secrets Use an approved secret mechanism where the environment and service policy permit it. Credential ownership, scope, rotation and the operation that the credential may perform. Never put raw tokens, keys, passwords or signing material in prompts, fixtures, logs or patches.
Sandbox Operate within the configured filesystem and network boundaries. Whether those boundaries are suitable for the repository and data involved. Prefer least privilege. Do not make unrestricted filesystem and network access the default.
Approvals Pause for approval when a command or operation falls outside the agreed routine. Who may approve and which actions remain prohibited even after a request. Sandbox reach and approval policy are different controls; require both authority and technical permission.
Release control Prepare a diff, check ledger, release-note candidate and unanswered decision list. Merge, version class, tag, signature, publication, announcement and backport decisions. Only an authorised human releases. No generated artefact counts as release approval.

OpenAI’s sandboxing guidance distinguishes technical sandbox boundaries from approval policy. It describes workspace-write with on-request as a lower-risk local automation preset. That is a useful starting point, not a universal configuration mandate. Managed-workspace rules and operating systems vary. Do not default to danger-full-access.

Codex CLI versus configured Codex Cloud

Choose the execution surface by where the authorised evidence and controls already reside. Codex CLI operates against the local repository and locally installed tools. The user controls those tools, permissions and the surrounding machine. It is appropriate when a maintainer has an approved checkout, can inspect every command and does not need Cloud-only infrastructure. Use a clean branch or worktree, record the baseline commit and check for pre-existing modifications before any edit.

Configured Codex Cloud uses prepared filesystem state. OpenAI’s Cloud environments documentation, accessed 2 October 2026, says each new task receives a separate workspace. That isolation is useful for avoiding accidental cross-task edits, but it does not itself prove that the setup is current, complete or authorised. Before relying on it, review the selected repository, setup process, dependency state and test result shown when the environment is prepared.

The same documentation says configured network secrets can be substituted for allowed HTTPS destinations. This does not grant service permission and is not a reason to request the secret in a prompt. Specify the required destination and operation abstractly—for example, “read public dependency metadata from the approved registry”—and stop if access is unavailable.

Use Codex CLI when local inspection and explicit command approval are the controlling requirements. Use Cloud only when the organisation has configured the repository, workspace, network and credentials for the task. If the reproduction requires an unapproved external service, prefer a local fixture or mock. If neither can represent the behaviour faithfully, record that limitation rather than manufacturing a result.

In both surfaces, separate concurrent work. A dedicated Git branch or worktree gives reviewers a known base and a bounded diff. It does not prevent every mistake, so begin with git status, record the base commit, and stop if unrelated changes are present. Do not reset, clean, overwrite or stash someone else’s work without explicit approval.

Why Chat, Work and API access are separate

OpenAI’s support article distinguishing Chat, Work and Codex, updated 2 October 2026, describes Codex as the product experience for software-development and technical work. Availability depends on plan, workspace settings and rollout. Do not infer that a normal ChatGPT conversation has terminal access, repository access or the same configured environment as Codex.

Likewise, access to ChatGPT or Codex does not establish access to the API, an API key, API billing or an API-hosted execution design. This playbook is not an API implementation guide. If a team wants programmatic orchestration, it must assess that architecture, credentials and service controls separately rather than copying these repository prompts into an unattended release system.

The decision rule is: name the surface actually being used and verify its permissions before the task. If the interface cannot inspect the authorised repository or execute the approved command, treat its output as advisory text only. Never imply tool execution from a chat response that did not execute tools.

The five-label prompt format and repository guardrails

OpenAI’s current Codex best-practices guidance recommends specifying a goal, context, constraints and a definition of done, and recommends planning, testing and reviewing changes. This article presents every maintenance prompt through five editorial labels: Purpose, Copy-paste prompt, Required inputs, Expected output and Verification checkpoint. The fixed five-label structure is this article’s device, not an OpenAI-mandated prompt syntax. Within the copyable text, role and outcome express the goal; authorised evidence supplies context; boundaries provide constraints; checks define done; and the requested artefact makes review practical.

Codex reads AGENTS.md instructions before work. OpenAI’s current guidance says project instructions are layered from the repository root towards the current directory, with nearer files applied later. It also documents a default 32 kibibytes combined-guidance limit. A root file might define public API policy and standard checks, while a nearer file might define parser-specific fixtures. The nearer instruction can specialise the root rule, but it should not silently overturn project governance.

Instruction files are evidence, not infallible policy. They may be stale, incomplete, contradictory or omitted because the combined guidance was truncated. Ask Codex to list discovered files, their scope and conflicts before relying on them. Compare consequential rules with canonical project documents such as the contribution guide, security policy and release policy. If instructions disagree about destructive commands, public compatibility or authority, choose the safer interpretation and stop for a maintainer.

Set stop conditions before investigation. Stop when the task would expose a secret or private report; modify files outside the approved path; contact an unapproved network service; destroy local state; change generated or vendored material without policy support; interpret a legal or security issue; or make a release decision. Uncertainty is not a defect to hide. Record it beside the claim it limits.

Prompt 1: Read the issue packet

Purpose

Convert an authorised issue packet into a traceable intake ledger without deciding that the report is valid. This first pass separates reporter statements from repository evidence and flags sensitive material before it spreads into logs or prompts.

Copy-paste prompt

You are assisting an authorised maintainer with issue intake. Read only the supplied issue packet and repository files explicitly listed below. Build a factual intake ledger for the fictional parsing-regression report.

Record the reported package version, prior comparison version, runtime, operating environment, input, expected result, actual result, reproduction steps and attached evidence. For every entry, distinguish a direct quotation or supplied fact from your interpretation. Do not diagnose the cause, edit files, run commands, browse external services or infer missing values.

Redact personal data, credentials, private URLs and unrelated customer content. Do not reproduce private vulnerability details. If the packet suggests a security issue, sensitive disclosure, licence dispute or missing maintainer authority, stop and identify the authorised escalation route if one is supplied.

Mark contradictions and unknowns explicitly. Produce only a reviewable intake ledger, a list of missing evidence and a recommendation to continue, request clarification or escalate. The recommendation is advisory; do not close, label or re-route the issue.

Required inputs

  • The authorised issue text or a sanitised export, including comments that are in scope.
  • Relevant attachment names and hashes or paths; exclude raw secrets and unnecessary personal data.
  • The repository’s public issue, conduct, support and security-routing policies, where applicable.
  • The maintainer’s statement of authority and the permitted data boundary.

Expected output

A table whose rows identify the claim, supplied value, source location, evidence type and uncertainty. Supported facts must be attributable to the packet. “The reporter states that fictional version 2.5.0 rejects a compact +0000 offset” is supported as a report; “2.5.0 contains a regression” is not yet established. Human decisions—such as whether to request more information, apply a project label or move to private handling—must appear separately.

Verification checkpoint

  • Confirm that no command was run and no repository change occurred.
  • Compare quoted values with the original packet.
  • Check that personal information, tokens and private vulnerability content were excluded.
  • Reject any causal conclusion not supported by a source location.
  • Require an authorised human to approve continued handling if security, privacy, legal, financial or conduct concerns appear.

A hypothetical intake ledger for Northstar software development kit (SDK)A collection of libraries, tools and documentation for building against a platform. Open glossary entry could begin as follows. It records reported information, not a verified outcome:

Item Ledger entry Evidence status Uncertainty or human action
Reported behaviour Reporter says fictional version 2.5.0 rejected parseTimestamp("2026-10-02T09:30:00+0000"). Supported only as a reporter statement in the fictional issue. Not reproduced; exact invocation, error and output capture required.
Comparison Reporter says fictional version 2.4.2 accepted the same compact-offset timestamp. Unverified comparison claim. Need equivalent environment and synthetic fixture for both versions.
Environment Runtime family supplied, patch version omitted. Partially supported by issue text. Ask for exact runtime and installation method.
Expected contract Reporter links to a parser example. Link presence is supported; contractual meaning is not assessed. Trace canonical documentation and exported types later.
Routing No secret or security allegation appears in the sanitised packet. Limited to supplied content. A maintainer still decides the project label and next step.

Prompt 2: Map applicable AGENTS.md instructions

Purpose

Identify which repository instructions govern the prospective investigation, how they layer, and whether any conflict or truncation risk requires human interpretation.

Copy-paste prompt

Inspect the authorised repository without editing it. Identify the AGENTS.md files that apply from the repository root to the directories likely involved in the parsing issue. Record each file’s path, scope and relevant instruction, preserving the order in which broader and nearer guidance applies.

Summarise only rules relevant to issue handling, parser code, fixtures, tests, supported versions, public API changes, generated files, networking, security routing and release authority. Do not treat instructions as automatically current or correct. Flag ambiguity, apparent conflict, missing canonical references and possible omission caused by the combined-guidance limit.

Cross-reference only the repository policy files I authorise. Do not resolve governance conflicts yourself. Do not edit AGENTS.md, install tools, run tests or access the network.

Return an instruction map that separates repository text from your interpretation. End with explicit stop conditions and questions requiring a maintainer decision.

Required inputs

  • The repository root and anticipated parser/test directories.
  • Read permission for applicable AGENTS.md and authorised policy files.
  • The issue’s likely file scope, marked as preliminary.
  • Any managed-workspace instructions that legitimately apply.

Expected output

An ordered map showing root-to-directory instruction layering, with source paths and concise quotations or faithful paraphrases. For example, a root requirement to run the public API test suite and a parser-directory requirement to use table-driven fixtures can both apply. If a nearer file appears to permit behaviour prohibited by the root governance policy, the output should flag a conflict rather than choosing whichever instruction is more convenient.

Verification checkpoint

  • Verify every mapped rule against its repository path.
  • Check whether the intended working directory changes which files apply.
  • Treat the documented default 32 kibibytes combined-guidance limit as a reason to inspect for omitted guidance, not as proof that truncation occurred.
  • Confirm that no instruction was edited and no conflict was silently resolved.
  • Have a maintainer decide conflicts involving security, compatibility, release policy or destructive commands.
  • If maintainer guidance or a rule file includes a sensitive personal detail or secret, redact it from the prompt and escalate to a human maintainer before proceeding.

Prompt 3: Classify scope and authority

Purpose

Classify the maintenance request before code execution. The important distinction is not merely “bug or not”; it is whether the supplied evidence belongs in public bug handling, support, documentation, feature discussion, contributor-policy review or private security escalation.

Copy-paste prompt

Using the reviewed intake ledger and instruction map, classify the report as one or more of: possible bug, feature request, documentation gap, support question, possible security report or contributor-policy issue.

For each classification, cite the supplied evidence that supports it and state what remains unknown. Assess only whether the authorised maintainer and current task have enough scope to continue with a non-destructive reproduction plan. Do not decide issue labels, disclosure status, fault, contributor intent, severity, compatibility policy or release priority.

Stop if the packet contains private vulnerability information, credentials, personal data beyond the approved need, an authority dispute or a request for exploit-oriented work. Do not transform sensitive details into a public summary.

Produce a routing recommendation with allowed next actions, prohibited actions, missing approvals and questions for the human maintainer.

Required inputs

  • The human-reviewed intake ledger.
  • The reviewed instruction map and applicable public policies.
  • A statement of repository, issue and command authority.
  • Any already-approved escalation route, referenced without confidential contents.

Expected output

A classification matrix that links each proposed category to evidence and uncertainty, followed by a bounded routing recommendation. For the fictional report, “possible bug requiring reproduction” may be reasonable if the claimed behaviour conflicts with a cited example. “Confirmed regression” would be premature. Human decisions must include public versus private handling, label selection, contributor response and permission to execute a reproduction.

Verification checkpoint

  • Confirm that classifications are phrased as provisional where evidence is incomplete.
  • Ensure a security signal causes a stop rather than a public debugging request.
  • Verify that authority covers the repository and proposed local operations.
  • Check that no judgement about a reporter’s motive or competence appears.
  • Require human approval before routing, contacting the reporter or continuing with consequential material.

Prompt 4: Design a minimal reproduction

Purpose

Design the smallest experiment capable of distinguishing the reported behaviour from a setup problem, without changing production code or reaching unnecessary services. A minimal reproduction controls variables; it does not attempt to prove every supported configuration.

Copy-paste prompt

Design, but do not execute, a minimal reproduction for the authorised Northstar SDK parsing report. Use only repository-provided tooling, local fixtures and the approved versions. Compare the reported current version and prior version under equivalent conditions where the repository supports that comparison.

Specify the exact fixture, entry point, commands, expected observations and environment fields to record. Keep network access disabled unless a maintainer separately approves a named destination and operation. Do not install unapproved packages, use real customer data, paste secrets, modify production code, publish artefacts or clean unrelated files.

Separate the reporter’s expected result from the repository-derived expectation. Include outcomes for reproduced, not reproduced, environment-dependent and blocked. State what the experiment cannot establish, including broad compatibility and root cause.

Return a reproduction plan, files that would be created, command approval requests and stop conditions. Wait for human approval before execution.

Required inputs

  • The accepted intake and routing decision.
  • Repository-provided setup and test commands.
  • Approved package/runtime versions and a synthetic fixture.
  • The isolated branch or worktree path and baseline commit.
  • The allowed filesystem and network boundary.

Expected output

A proposed experiment with one variable changed at a time. An example is to invoke the same public parser with the same synthetic compact-offset timestamp fixture under versions 2.5.0 and 2.4.2, while recording exact runtime and dependency resolution. This is an illustrative design, not an executed test. The plan should explain that differing outputs would support a version-associated behavioural change, but would not by itself identify the responsible commit or prove that the newer behaviour violates the contract.

Verification checkpoint

  • Review every command before execution and reject destructive or unexplained flags.
  • Ensure fixtures contain no customer records, credentials or private report content.
  • Confirm the comparison keeps runtime, invocation and input equivalent.
  • Check that expected observations came from an authorised source rather than model inference.
  • An authorised maintainer must approve the commands, created files and any network access.
  • When the reproduction is inconclusive or a required dependency state is unknown, mark it unresolved and ask a human maintainer for a bounded next test rather than inventing a result.

Prompt 5: Execute only approved reproduction steps

Purpose

Run the approved, non-destructive experiment and preserve exact observations. This step must not drift into patching, dependency upgrades or root-cause speculation.

Copy-paste prompt

Execute only the reproduction steps explicitly approved below in the named isolated branch or worktree. First report the current directory, baseline commit and working-tree status. Stop if the baseline differs, unrelated changes exist, a required tool is missing, a command requests broader access or an unapproved network connection would occur.

For each approved command, record the exact command, exit status, relevant standard output and standard error, runtime and package versions, fixture hash and files created or changed. Preserve errors rather than retrying with broader permissions. Do not edit source code, upgrade dependencies, clean files, publish results externally or use secrets.

Classify the observation only as reproduced, not reproduced, environment-dependent or blocked according to the approved criteria. Do not claim correctness, compatibility, security or root cause.

Return a reproduction record and a list of deviations or uncertainties. Stop after the approved commands and wait for human review.

Required inputs

  • The human-approved reproduction plan and exact command list.
  • The approved isolated worktree or branch and expected baseline commit.
  • The synthetic fixture and its expected hash.
  • Approved environment versions and explicit network policy.

Expected output

A chronological record, including blocked or failed commands. The following is a hypothetical format, not a claimed execution result:

Record field Example entry Interpretation limit
Baseline <approved commit identifier>; clean worktree reported before execution. Must be checked by the maintainer; placeholder is not evidence.
Fixture 2026-10-02T09:30:00+0000 stored in an approved temporary test path. Synthetic compact-offset input covers one illustrative case only.
Current-version command <repository-approved parser command> No result is asserted until actually run and captured.
Prior-version command <equivalent approved comparison command> Comparison is invalid if dependency or runtime conditions differ materially.
Observed outputs <insert verbatim captured outputs after execution> Do not prefill with the reporter’s expected values.
Classification <reproduced | not reproduced | environment-dependent | blocked> A human checks the classification against the approved criteria.
Residual uncertainty Other input forms, runtimes and supported versions not tested. No broad compatibility conclusion follows.

Verification checkpoint

  • Compare executed commands character-for-character with the approved plan.
  • Inspect git status and the created-file list without deleting evidence.
  • Check outputs against raw command logs, not a generated summary alone.
  • Treat a passing or failing command as one observation, not proof of release safety.
  • Have the maintainer decide whether evidence is sufficient for contract analysis or whether a revised experiment needs fresh approval.

Prompt 6: Trace the public contract

Purpose

Map the observed behaviour to the package’s public promises before proposing a change. Public contract evidence can include exported symbols, documented input and output types, examples, error behaviour and supported-version policy. Internal implementation habits are not automatically public guarantees.

Copy-paste prompt

Using the human-reviewed reproduction record, trace the public contract relevant to the parsing behaviour. Inspect only authorised repository sources: exported parser symbols, public type declarations, API reference text, README examples, changelog entries, migration guides, tests explicitly presented as contract fixtures and supported-version policy.

For each source, quote or precisely reference the path and explain whether it states a requirement, demonstrates an example, records historical behaviour or merely reflects implementation. Distinguish normative language from inference. Search for contradictory promises and version-specific qualifications.

Do not decide semantic versioning, deprecation, whether undocumented behaviour must be preserved, or whether the reporter is entitled to a particular result. Do not edit code or documentation. Do not use external material unless a maintainer separately authorises it.

Return a contract-impact map covering the input, output, error behaviour, affected public entry points, evidence strength, contradictions and unresolved maintainer decisions.

Required inputs

  • The reviewed reproduction record, including raw observations.
  • Authorised public API declarations, documentation, examples and release history.
  • The project’s support and compatibility policy.
  • A clear distinction between generated files and their canonical sources.

Expected output

A source-indexed map, not a verdict. For example, an exported parameter type accepting a string would be relevant but may not establish whether a compact numeric UTC offset is valid. A README example showing +0000 is useful evidence but may be illustrative rather than exhaustive. A regression test added for a past issue may establish historical intent, subject to maintainer interpretation. The artefact should expose these distinctions instead of collapsing them into “breaking” or “safe”.

Verification checkpoint

  • Open every cited repository path and verify the quoted wording and version context.
  • Confirm that generated documentation is traced to its canonical source where possible.
  • Reject conclusions based solely on implementation details or one generated test.
  • List contradictions between types, examples, tests and policy rather than selecting a preferred source silently.
  • Reserve compatibility, deprecation, semantic-version and release decisions for authorised maintainers.

At the end of this phase, the project should possess reviewable intake, instruction, authority, reproduction and contract records. It should not yet possess an assumed fix or release conclusion. Continue only if a human accepts the evidence boundary, confirms the work remains authorised and decides that deeper compatibility analysis is warranted.

The contract-impact map from Prompt 6 should exist before anyone proposes an edit. That order matters because an open-source package is consumed through more than its implementation: published documentation, exported functions, CLI flags, type declarations, examples, error behaviour and declared runtime support can all form part of users’ expectations. A locally plausible code change may therefore repair one report while contradicting a documented contract or dropping a supported runtime.

The running example in this section is fictional. “Northstar SDK” is an illustrative package whose parser accepts timestamps returned by an authorised service. A fictional issue claims that version 2.5.0 rejects parseTimestamp("2026-10-02T09:30:00+0000"), while version 2.4.2 accepted the same compact UTC-offset input. No outcome, test or report shown below was observed in a real repository.

A contract-impact map inherited from Prompt 6 might contain the following illustrative entries:

Contract surface Repository evidence to locate Possible impact Open question for a maintainer
Exported parser src/parseTimestamp.ts and package export declarations Changing accepted input could affect every caller of the public function. Is the accepted timestamp grammar documented or intentionally delegated to a runtime parser?
Type declaration Public parameter and return types The type may promise only a string, while prose documentation imposes a narrower format. Which source defines the public constraint when the type and prose differ?
Usage guide README and reference examples An example containing +0000 would strengthen the case that rejection is a regression. Is the example normative policy, historical illustration or obsolete documentation?
Error behaviour Tests and documented exceptions Accepting another format may change which inputs throw an error. Does the project promise a particular error class or message?
Supported runtimes Package metadata, test matrix and support policy A replacement parsing method might not exist or behave consistently on every declared runtime. Which declarations are authoritative, and are they consistent?
Downstream compatibility Repository fixtures and authorised issue reports Normalising input could preserve successful parsing but alter the returned representation. Do consumers observe the normalised value, the instant, or both?

Assess declared support before editing. First identify what the repository says it supports; then determine whether the reported environment falls inside that statement; only then consider implementation choices. A failing example on an unsupported runtime can still justify documentation, support or future-feature work, but it is not automatically evidence of a regression within the maintained contract. Conversely, a failure on a declared runtime cannot be dismissed merely because a newer runtime behaves differently.

Repository statements can conflict. Package metadata may set one minimum runtime, continuous-integration configuration another and prose documentation a third. Treat each as evidence, not permission to choose whichever makes the patch easiest. Record the conflict, identify any documented precedence and ask an authorised maintainer to resolve it. Generated analysis must not silently invent a support policy or infer semantic versioning treatment from the size of a diff.

Four figures in parallel lanes passing a component, test blocks, and evidence tokens between bounded work areas
Separate lanes clarify that the maintainer, Codex, existing repository checks, and contributors have different responsibilities and cannot silently substitute for one another.

Decision table: choose a response class before choosing code

This table separates response classes; it does not assign a version number. Semantic versioning, backporting and deprecation treatment depend on the project’s own policy and authorised maintainer judgement. A one-line edit can alter a major public contract, while a larger internal refactor can preserve observable behaviour.

Response class Use when evidence supports Main trade-off Evidence required before selection Human decision gate
No change The report is not reproducible within declared support, expected behaviour is already explicit, or the requested behaviour conflicts with an intentional contract. Avoids unnecessary churn but risks dismissing a genuine edge case if evidence is incomplete. Reproduction record, applicable policy, supported-version finding and a clear account of unresolved uncertainty. A maintainer decides whether evidence is sufficient to close, defer or request more information.
Support response The behaviour is expected and the reporter needs correct usage, configuration or an existing migration route. Fast and low-impact, but inadequate if the instructions themselves are missing or misleading. Verified usage path and an authoritative repository source supporting the response. A maintainer approves the public reply and any request for additional non-sensitive diagnostics.
Documentation clarification Implementation matches intended policy, but public wording, examples or error guidance are ambiguous. Clarifies expectations without runtime change, but cannot repair behaviour that violates the actual contract. Confirmed implementation behaviour, intended policy and exact documentation locations. A maintainer confirms that the wording reflects policy rather than creating it.
Patch A bounded implementation defect contradicts the supported contract and can be corrected without an approved contract change. Restores intended behaviour, but may expose untested interactions or alter observable errors and normalisation. Reproduction, contract evidence, affected-version analysis, scoped plan and regression-test design. A maintainer approves scope, implementation approach and later release classification.
Deprecation Existing behaviour must remain temporarily available while users move to a replacement. Reduces abrupt disruption but adds parallel behaviour, communication work and a future removal decision. Current contract, replacement path, warning surface, migration feasibility and project deprecation policy. Authorised humans set wording, timing, support period and any eventual removal.
Breaking change The desired correction necessarily changes an intentional public contract or removes supported behaviour. May simplify or correct the package, but transfers migration work and risk to downstream users. Demonstrated incompatibility, alternatives considered, migration impact and relevant release policy. Maintainers decide whether, when and under which release class the change can proceed.

A practical decision rule is to select the least disruptive response that satisfies verified policy and the confirmed user need. “Least disruptive” does not mean “smallest diff”: it means the narrowest change to the public contract. If authority, policy or evidence is missing, stop at an option comparison rather than converting uncertainty into code.

Prompt 7: Check supported versions and compatibility policy

Purpose

Determine whether the reported package version, runtime, platform and dependency combination falls within declared support. This prompt searches repository-controlled material; it does not let Codex invent a compatibility policy. Run it before changing implementation code because an apparently attractive API or runtime feature may be unavailable on supported versions.

Copy-paste prompt

Goal: Build an evidence-based support and compatibility assessment for the parsing report, without editing files.

Context:
- Repository root: [PATH]
- Contract-impact map from Prompt 6: [PATH OR PASTE]
- Reported package version: [VERSION]
- Reported runtime/platform/dependencies: [DETAILS]
- Authorised issue or reproduction record: [PATHS]
- Candidate policy locations, if known: [PATHS]

Constraints and stop conditions:
- Read applicable repository instructions, including active AGENTS.md files, before analysis.
- Search only this authorised repository and supplied records.
- Never paste, request or expose secrets, private vulnerability details or unrelated user data.
- Distinguish direct repository evidence from inference in every conclusion.
- Do not infer support merely from code compiling, a dependency installing or a test job existing.
- If policies conflict, appear stale, are truncated, or no source establishes authority, report that and stop short of choosing a policy.
- If access, credentials or maintainer authority are missing, stop and list what an authorised human must provide.
- Do not assign semantic versioning, deprecation treatment or release eligibility.
- Await human approval before any edit or external query.

Required checks:
1. Inspect package metadata, support-policy files, current documentation, test matrices and relevant history available locally.
2. Record exact paths and quoted or tightly paraphrased declarations.
3. Compare the report’s package, runtime, platform and dependency versions with each declaration.
4. Identify conflicts, gaps and ambiguous terms such as “current”, “recent” or “supported”.
5. Note whether a proposed implementation technique would exclude any declared environment.

Done when:
Return an affected-version table with columns for dimension, reported value, declared support evidence, assessment, confidence and uncertainty. Follow it with:
- repository facts;
- clearly labelled inferences;
- unresolved human decisions;
- an explicit recommendation to continue, request evidence or stop.
Do not modify the working tree.

Required inputs

  • The Prompt 6 contract-impact map and Prompt 5 reproduction record.
  • The reported package, runtime, operating-system and relevant dependency versions.
  • Repository policy files, package metadata and test configuration available under the maintainer’s authorised access.
  • Any documented precedence rule for conflicts between metadata, documentation and automation.

Expected output

The central artefact is an affected-version table. The following is illustrative, not an observed Northstar SDK result:

Dimension Reported value Declared support evidence Assessment Confidence Uncertainty
Package 2.5.0 In this illustration, CHANGELOG.md identifies version 2.5.0 as current. Inside the current line. High No statement yet establishes whether older 2.x lines receive fixes.
Runtime Runtime 18 Package metadata says >=18; a prose guide says “20 or later”. Conflicting declarations; no conclusion. High confidence in conflict An authorised maintainer must identify the controlling policy.
Operating system Linux Test configuration lists Linux jobs but no prose support statement. Test coverage is evidence of testing, not by itself a support promise. Medium Platform policy is absent or not found.
Timestamp form +0000 offset README example uses the form; reference grammar is silent. Possible public expectation, not yet confirmed as normative. Medium Maintainer interpretation is required.

Verification checkpoint

  • Supported source facts: paths, declarations and configuration actually present in the authorised repository.
  • Documentation basis: the current OpenAI best-practices guidance supports giving Codex a goal, context, constraints and a done condition.
  • Human decisions: which conflicting declaration controls, whether an example is contractual, and which versions should receive a correction.
  • Verify every quotation against the file and revision cited. A test matrix must not be promoted automatically into a support guarantee.
  • If authority is missing, credentials are required or policy is contradictory, stop, report uncertainty and await maintainer approval.

Prompt 8: Search authorised regression evidence and duplicates

Purpose

Find whether the behaviour was introduced by a repository change, is already covered by an issue or test, or has a competing explanation. Search only sources the maintainer is authorised to inspect. Absence of a matching issue is not proof that no duplicate exists, and textual similarity is not proof that two reports share a cause.

Copy-paste prompt

Goal: Search authorised local evidence for duplicates, regression signals and competing explanations for the timestamp parsing report.

Context:
- Issue packet: [PATH]
- Reproduction record: [PATH]
- Contract and support assessments: [PATHS]
- Authorised local history range: [RANGE]
- Search terms or symbols: [TERMS]

Constraints and stop conditions:
- Read applicable AGENTS.md guidance first.
- Use only the local repository, supplied issue exports and other explicitly authorised records.
- Do not query external issue trackers, services or private systems without separate human approval and credentials already configured through approved means.
- Never request raw tokens, private vulnerability material or unrelated logs.
- Label each statement as repository evidence, supplied-report evidence or inference.
- Do not claim causation from timing, matching text or a single changed line.
- If history is shallow, issue data is incomplete, or access is denied, report the limitation.
- Stop on missing authority or credentials and await human approval before expanding the search.

Required checks:
1. Search tests, fixtures, changelog entries, issues and commit messages for the input form and affected symbol.
2. Inspect relevant blame/history only within the authorised range.
3. Compare candidate duplicates by symptom, environment, affected API and reproduction.
4. Identify commits that changed parsing, validation, normalisation or runtime support.
5. Produce multiple hypotheses, including non-regression explanations where evidence permits.

Done when:
Return:
- an evidence index with paths or authorised identifiers;
- candidate duplicates with match and mismatch fields;
- a confidence-ranked hypothesis list;
- missing evidence and proposed next checks;
- a stop/continue recommendation requiring maintainer approval.
Make no edits.

Required inputs

  • Authorised issue text stripped of unnecessary personal or confidential information.
  • The exact failing input and observed error from the approved reproduction record.
  • A local history range, issue export or connected source the maintainer is permitted to search.
  • The affected public symbol and likely implementation paths.

Expected output

An illustrative confidence-ranked hypothesis list might read as follows; it is not a finding about a real package:

  1. Medium confidence — parser dependency behaviour changed during the fictional 2.5.0 update. Repository evidence: a dependency range changed in package.json and parsing tests cover +00:00 but not +0000. Inference: the dependency change may explain the symptom. Missing evidence: before-and-after execution under the approved matrix.
  2. Medium-low confidence — a validation expression became narrower. Repository evidence: local history shows an expression changed near the parser. Inference: the expression could reject compact offsets. Missing evidence: confirmation that the changed path handles the reported public API.
  3. Low confidence — runtime differences account for the report. Repository evidence: support declarations conflict and the reproduction covers only one runtime. Inference: another runtime may parse differently. Missing evidence: approved comparison on every relevant supported runtime.
  4. Low confidence — the report duplicates fictional earlier issue N-142. Supplied issue evidence: both mention timestamp rejection. Mismatch: N-142 concerns a missing timezone rather than a compact offset. Decision: do not merge or close as a duplicate without human review.

Confidence ranks prioritise investigation; they do not turn a hypothesis into a diagnosis. The useful distinction is between “the repository contains this change” and “this change caused the report”. The latter requires targeted evidence.

Verification checkpoint

  • Supported source facts: exact commits, paths, tests and authorised issue identifiers found during the bounded search.
  • Human decisions: whether reports are duplicates, whether more history may be accessed and which hypothesis deserves investigation.
  • Check that every causal statement is labelled as inference until a discriminating test supports it.
  • On shallow history, unavailable credentials, private security content or uncertain search authority, stop and await an authorised maintainer.

Prompt 9: Compare response options and trade-offs

Purpose

Convert the contract and regression evidence into alternatives rather than rushing to a favoured fix. This is the gate where no change, support guidance, documentation, a patch, deprecation and a breaking change are compared against the same evidence. Codex prepares the matrix; maintainers choose the project response and later determine release treatment.

Copy-paste prompt

Goal: Compare viable response classes for the authorised parsing issue without selecting one as project policy.

Context:
- Contract-impact map: [PATH]
- Support assessment: [PATH]
- Regression and duplicate evidence: [PATH]
- Repository compatibility/deprecation policy: [PATH OR “not established”]
- Confirmed reproduction: [PATH OR “not reproduced”]

Constraints and stop conditions:
- Distinguish repository evidence from inference and proposal.
- Compare no change, support response, documentation clarification, patch, deprecation and breaking change; mark genuinely inapplicable options with evidence.
- Do not assign semantic versioning, release timing, backports or deprecation duration.
- Do not treat implementation effort as more important than public-contract impact.
- Do not assume a passing test would prove compatibility or release safety.
- If policy authority, credentials or essential evidence is missing, stop at a conditional matrix.
- Report uncertainty and await human approval before planning an implementation.
- Keep secrets and private vulnerability details out of the prompt and output.

Required checks:
1. State what each option preserves and changes for users.
2. Identify documentation, test and migration obligations.
3. Identify risks, reversibility and unresolved evidence.
4. Separate a behavioural correction from a contract change.
5. Name the exact maintainer decisions required.

Done when:
Return an option matrix, a short evidence-bound comparison, disqualifying conditions for each option and numbered human decision questions. Do not recommend an option unless the evidence clearly eliminates all others; even then, label it a proposal awaiting approval.

Required inputs

  • The outputs of Prompts 6–8.
  • Any explicit compatibility, deprecation and contribution policies.
  • Known consumer-visible effects: accepted input, returned value, errors, types and documentation.

Expected output

The following option matrix is illustrative rather than observed:

Option Contract effect Potential advantage Risk or cost Evidence gap Decision owner
No change Preserves rejection of compact offsets. Avoids broadening accepted input without authority. Could leave a documented example unusable. Whether the example is normative. Maintainer
Support response Asks users to convert +0000 to +00:00. No runtime change. Moves conversion burden downstream and may conflict with service output. Whether users control the source format. Maintainer or support owner
Documentation clarification States the currently accepted grammar. Resolves ambiguity if rejection is intentional. Would conceal rather than repair a regression if prior acceptance was promised. Intended grammar. Maintainer and documentation reviewer
Patch Accepts both compact and colon-separated UTC offsets. Could restore intended acceptance narrowly. Normalisation may alter returned text or permit unintended forms. Cross-runtime behaviour and negative cases. Maintainer
Deprecation Temporarily preserves one path while announcing a replacement. Provides migration time if grammar must narrow later. Adds warning and migration complexity. Whether a replacement API exists. Project governance
Breaking change Defines and enforces a new timestamp grammar. Could make behaviour explicit and consistent. Requires downstream migration and may be disproportionate. Evidence that current contract cannot be preserved. Authorised release maintainers

Verification checkpoint

  • Supported source facts: documented behaviour, confirmed reproduction, applicable policy and observed repository history.
  • Human decisions: intended contract, acceptable disruption, response class, deprecation route and eventual release classification.
  • Reject any matrix that equates line count with compatibility impact or labels an option “safe” without qualifications.
  • If the controlling policy or decision owner is unknown, report uncertainty, stop and await human approval.

Prompt 10: Write a scoped implementation plan

Purpose

Translate the selected response into a reviewable, file-level plan. Planning is appropriate only after a human chooses the response class. The plan should name non-goals and revert conditions so a narrow parsing correction does not become an unapproved parser redesign, dependency migration or support-policy change.

Copy-paste prompt

Goal: Draft a file-level implementation plan for the maintainer-approved option; do not edit code.

Approved response and boundaries:
- Human-approved option: [OPTION]
- Approval record: [REFERENCE]
- Contract that must be preserved: [SUMMARY]
- Allowed behaviour change: [SUMMARY]
- Explicit non-goals: [LIST]
- Authorised repository and base revision: [DETAILS]

Constraints and stop conditions:
- Read active AGENTS.md instructions and report conflicts or possible truncation.
- Separate repository facts, engineering inference and human-approved decisions.
- Do not broaden support, change public types, add dependencies, alter generated files or rewrite unrelated code unless explicitly approved.
- Do not choose semantic versioning, deprecation timing, backports, merge or publication actions.
- Never request secrets or private vulnerability details.
- If a required file is generated, ownership is unclear, credentials are needed or the approved option cannot preserve its stated contract, stop.
- Report uncertainty and await human approval before implementation.

Required checks:
1. Identify each file expected to change and why.
2. Identify test files and specify the causal regression claim.
3. List files inspected but intentionally unchanged.
4. Describe edge cases, negative cases and supported-version concerns.
5. Define deviation, abort and revert conditions.
6. List reviewer questions that cannot be resolved from repository evidence.

Done when:
Return a numbered file-level plan, non-goals, proposed validation commands taken from repository instructions, rollback/revert conditions, uncertainties and an approval checklist. Make no changes.

Required inputs

  • A recorded maintainer selection from Prompt 9.
  • The approved public-contract boundary and affected support matrix.
  • Applicable repository guidance and prescribed validation commands.
  • Locations of implementation, tests, documentation and generated outputs.

Expected output

An illustrative file-level plan could be:

  1. src/parseTimestamp.ts: add a narrowly scoped normalisation step for the approved compact UTC-offset form before the existing parse call. Preserve the current return type and existing handling of colon-separated offsets.
  2. test/parseTimestamp.test.ts: add one regression fixture for the approved form, one existing accepted form and negative cases that must remain rejected. State that the test demonstrates these examples only, not all timestamp compatibility.
  3. README.md: no change under the current approval because the existing example already expresses the intended accepted form. Flag any ambiguous grammar prose for a separate documentation decision.
  4. package.json: inspect prescribed commands but do not change dependencies, engines or scripts.

Illustrative non-goals: replacing the parser library; defining a complete timestamp standard; changing error classes; broadening runtime support; backporting; assigning a release number.

Illustrative abort conditions: the normalisation changes the returned public representation; the repository’s supported runtimes implement the proposed operation differently; generated-source rules require an unavailable tool; or a negative case becomes accepted. Each condition returns the work to a maintainer rather than inviting an improvised workaround.

Verification checkpoint

  • Supported source facts: actual file ownership, scripts, generation rules and active repository instructions. OpenAI’s current guidance says Codex reads layered AGENTS.md guidance, with nearer files applied later and a default combined limit of 32 kibibytes; that guidance may still be stale, conflicting, incomplete or truncated.
  • Human decisions: approved response, non-goals, acceptable observable changes and whether documentation belongs in this patch.
  • Compare the plan line by line with the approval record. Any extra dependency, public type change or policy edit requires renewed approval.
  • Stop on missing authority, inaccessible tools or unresolved instruction conflicts; report uncertainty and await approval.

Prompt 11: Establish an isolated work context and baseline

Purpose

Create a reviewable starting point before an edit. Local Codex CLI execution uses the reader’s installed tools, filesystem access and permissions. According to OpenAI’s CLI page accessed on 2 October 2026, Codex CLI can inspect, edit and run local code, and can support repeatable automation; those capabilities do not grant repository rights or release authority.

A prepared Codex Cloud environment is also not an authorisation shortcut. OpenAI’s Cloud-environment documentation, accessed on 2 October 2026, says published environments reuse prepared filesystem state and each new task receives a separate workspace. Configured network secrets may be substituted for allowed HTTPS destinations, but workspace preparation does not grant repository access, service permissions or permission to contact an external destination. Never put raw credentials in a prompt.

Copy-paste prompt

Goal: Verify an isolated, reviewable work context and record the baseline before any patch.

Context:
- Approved plan: [PATH]
- Repository root: [PATH]
- Expected base branch/revision: [VALUE]
- Intended branch or worktree: [VALUE]
- Allowed files: [LIST]
- Prescribed baseline commands: [COMMANDS]

Constraints and stop conditions:
- Inspect status before changing anything.
- Do not clean, reset, stash, switch branches, discard files or overwrite uncommitted work.
- Do not create remote branches, contact external services or use credentials without explicit approval.
- Keep repository evidence separate from inference about who owns existing changes.
- Treat sandbox boundaries and approval policy as separate controls.
- Do not request danger-full-access; use the least access compatible with the approved task.
- Stop if the tree is dirty, the base differs, nested repositories are unexpected, credentials are missing or isolation cannot be confirmed.
- Report uncertainty and await human approval before setup changes or edits.

Required checks:
1. Record repository root, current branch, HEAD/base revision and status.
2. Identify untracked, modified, staged and ignored task-relevant files without exposing secret contents.
3. Confirm the intended worktree/branch and allowed-file boundary.
4. Record applicable instructions and available prescribed tools.
5. Run only approved non-destructive baseline commands.
6. Distinguish command output from conclusions.

Done when:
Return a baseline record containing revision, branch/worktree, status, applicable guidance, command results, allowed files and a prominent dirty-tree warning where needed. Do not edit implementation files.

Required inputs

  • The approved plan and expected base revision.
  • The intended local worktree, branch or separately configured Cloud task workspace.
  • Approved, non-destructive baseline commands.
  • The allowed file list and any repository-specific checkpoint convention.

Expected output

An illustrative dirty-tree warning should be explicit rather than buried:

OpenAI’s sandbox guidance distinguishes technical boundaries—what files and network resources commands can reach—from approval policy—when Codex must pause. It describes workspace-write with on-request as a lower-risk local automation preset. This does not override managed-workspace policy. Do not default to danger-full-access.

Verification checkpoint

  • Supported source facts: the recorded revision, status output, workspace location, available tools and applicable repository instructions.
  • Human decisions: ownership and disposition of pre-existing changes, acceptable isolation method and permission for any network or remote action.
  • Have a human compare the baseline with the intended base and confirm that no existing work will be overwritten.
  • If the worktree is dirty, credentials are unavailable or the workspace boundary is uncertain, stop and await approval. A separate Cloud workspace does not itself resolve repository authority.

Prompt 12: Make the smallest approved patch

Purpose

Implement only the approved plan and return a patch candidate for review. “Smallest” means the narrowest approved semantic change, not merely the fewest altered lines. A compressed expression that obscures edge cases can be less reviewable than a slightly longer, explicit check.

Copy-paste prompt

Goal: Implement the smallest patch that satisfies the approved plan, then report the semantic change for human review.

Context:
- Approved plan and approval reference: [PATHS]
- Clean baseline record: [PATH]
- Allowed files: [LIST]
- Required preserved behaviours: [LIST]
- Approved changed behaviour: [LIST]
- Prescribed formatting or focused-check commands: [COMMANDS]

Constraints and stop conditions:
- Re-read applicable AGENTS.md guidance before editing.
- Modify only allowed files and only for the approved behaviour.
- Do not add dependencies, change public types, revise support policy, alter generated files or perform unrelated cleanup without approval.
- Do not merge, commit, push, tag, publish, choose semantic versioning or claim the issue fixed.
- Never access external services or secrets unless separately authorised and configured; never print credentials.
- Distinguish repository evidence, implementation inference and human decisions.
- If implementation requires deviation, a broader contract change, unavailable credentials or edits outside scope, stop before making that deviation.
- Report all uncertainty and await human approval at the patch gate.

Required checks:
1. Inspect the final diff against the recorded baseline.
2. Confirm every changed hunk maps to an approved plan item.
3. Identify observable behaviour added, preserved and potentially altered.
4. Run only approved focused checks; record exact commands and results without treating them as proof.
5. Report skipped checks, limitations and any unexpected file changes.
6. Leave the result as an unapproved patch candidate.

Done when:
Return:
- changed files;
- a semantic diff summary;
- exact commands and outcomes;
- deviations (or “none observed”);
- residual uncertainties;
- reviewer questions;
- an explicit “awaiting human approval” status.
Do not self-approve the patch.

Required inputs

  • The human-approved Prompt 10 plan.
  • A clean or explicitly accepted Prompt 11 baseline.
  • The allowed-file boundary and preserved behaviours.
  • Approved focused checks that can run with the current tools and permissions.

Expected output

A semantic diff explains user-visible meaning rather than restating line additions. The following is illustrative, not an observed patch or test result:

Changed files: src/parseTimestamp.ts and test/parseTimestamp.test.ts.

Behaviour added: before invoking the existing parser, the candidate recognises only an otherwise valid timestamp ending in a four-digit numeric UTC offset and inserts the approved colon separator.

Behaviour intended to remain unchanged: colon-separated offsets continue through the existing path; timestamps without an offset retain their existing handling; public function names, parameter types, return types and documented errors are not intentionally changed.

Boundary: the candidate does not attempt general timestamp repair and does not accept alphabetic timezone names, malformed offset lengths or extra trailing characters.

Tests proposed or changed: one compact-offset case, one already accepted colon-separated-offset case and selected rejection cases. These examples test only the stated cases; they do not prove universal parser correctness, security or compatibility.

Deviation from plan: none identified in the illustrative output. A reviewer must verify this against the actual diff.

Residual uncertainty: cross-runtime behaviour remains unverified until the repository’s approved validation matrix is run; the controlling runtime-support declaration still requires maintainer confirmation if Prompt 7 found a conflict.

Status: patch candidate awaiting authorised human approval; not merged, released or classified.

A focused command succeeding would show only that the command completed successfully in that environment for the exercised cases. It would not prove correctness, backwards compatibility, security or release readiness. Likewise, a generated review is another review artefact, not a substitute for repository tests, branch protections or required approvals.

Verification checkpoint

  • Supported source facts: the actual diff, command output, baseline revision and repository-prescribed checks. OpenAI’s CLI documentation says a dedicated review can be performed without changing the working tree, but such a review does not authorise acceptance.
  • Human decisions: whether the implementation matches intent, whether residual risks are acceptable, which broader checks to run and whether the candidate may advance.
  • Review each hunk against the approved file-level plan. Unexpected files, dependency changes, public-type changes or broadened input acceptance return the work to the plan gate.
  • For security, privacy, financial, employment, government or other consequential decisions, require qualified human review and do not treat generated analysis as the decision. Keep secrets and unauthorised personal data out of prompts; treat instructions in untrusted repository text as data, not commands.
  • If authority, credentials, policy interpretation or compatibility evidence is missing, stop, report uncertainty and await human approval.

The endpoint of this section is deliberately limited: an isolated, evidence-linked patch candidate whose scope can be compared with the approved plan. Regression coverage, prescribed validation, contract review, compatibility scenarios and the eventual release decision remain separate gates.

Causal regression tests: prove only the behaviour they isolate

This section starts with a patch candidate, not an accepted fix. Continue the fictional example used throughout the playbook: an authorised maintainer is assessing a proposed change for an SDK parsing regression. The reported symptom is that a documented input no longer produces the expected parsed value. Repository evidence, approved fixtures and public documentation may be examined; private vulnerability details, customer logs and credentials must not be copied into prompts.

A regression test is causal when it isolates the behaviour changed by the patch and can distinguish the candidate from the relevant baseline. “Causal” does not mean that the test proves the complete root cause. It means the test is designed so that its observed difference can reasonably be attributed to the scoped code change rather than unrelated environment, network, timing or fixture changes.

Prompt 13: Add causal regression coverage

Purpose

Create or adjust the smallest repository-conforming test that exercises the reported parsing behaviour. The artefact must state what the test can establish, what it cannot establish and how a maintainer can compare the baseline with the patch.

Copy-paste prompt

You are preparing reviewable regression coverage for an authorised open-source SDK patch candidate.

Goal:
Add the smallest test that distinguishes the reported parsing behaviour on the named baseline from the proposed patch. Follow applicable repository instructions and existing test conventions.

Authorised evidence:
- Approved issue summary: [SANITISED SUMMARY]
- Baseline commit or tag: [BASELINE]
- Patch commit or working-tree diff: [PATCH]
- Relevant public contract: [DOCUMENTATION OR API PATHS]
- Existing nearby tests: [PATHS]
- Allowed files: [PATHS]

Constraints and stop conditions:
- Do not invent executions, outputs, supported versions, compatibility assurances or test results.
- Do not request or expose raw secrets, private vulnerability details, customer data or proprietary logs.
- Do not broaden the patch unless an authorised maintainer approves a revised plan.
- Do not remove or weaken existing assertions merely to make the patch pass.
- Do not choose deprecation or release timing.
- Do not approve your own patch or describe this test as proof of correctness, security, compatibility or release safety.
- If the behaviour cannot be isolated with the supplied evidence, stop and identify the missing fixture, contract statement or environment detail.
- For security-sensitive material, keep the task defensive and within the authorised scope; do not assume that this framing removes platform safety checks.

Required checks:
1. Read the applicable AGENTS.md instructions and report conflicts, gaps or possible truncation.
2. Identify the public behaviour represented by each new assertion.
3. Use a deterministic, repository-local fixture where the repository permits it.
4. Specify the exact command that a human may use on the baseline and patch.
5. Separate observed results from expected results. If you did not run a command, write “not run”.
6. Check that the test would not pass for an unrelated reason, such as asserting only that no exception occurred.
7. List adjacent cases deliberately excluded.

Fixed output artefact:
A regression-test proposal containing changed paths, test name, fixture rationale, assertion-to-contract mapping, exact comparison commands, observed or not-run status, expected baseline/patch distinction, limitations, and questions requiring human decision.

Required inputs

  • Supported source facts: the Codex best-practices guide recommends supplying a goal, context, constraints and a definition of done, and recommends testing and reviewing changes. The Codex CLI can inspect, edit and run local code, subject to user-controlled tools and permissions.
  • Repository evidence: the approved issue summary, baseline reference, patch diff, applicable public contract, nearby test conventions and allowed paths.
  • Human decisions: whether modifying production code is allowed, whether a new fixture is acceptable, which baseline represents the regression boundary and whether broader coverage is proportionate.

Expected output

The output should be a patch plus a bounded claim, not a claim that “the bug is fixed”. A worked example of the language change is:

Form Example claim Editorial decision
Before “The new test proves parsing is fixed and backwards compatible.” Reject. One test cannot establish all parsing behaviour or compatibility across untested consumers and versions.
After “Example proposed claim: for the supplied documented fixture, this assertion is expected to distinguish baseline [BASELINE] from patch [PATCH]. The commands have not yet been run. Other input forms, supported runtimes and downstream wrappers remain unverified.” Accept as a test proposal because it separates expectation from observation and names its exclusions.

A suitable test should normally assert the parsed value or documented error contract, not merely process completion. If the public promise concerns preservation of an empty field, for example, assert the exact representation the documentation specifies. Do not add an invented expected value where the contract is silent; mark that ambiguity for a maintainer.

Verification checkpoint

  • Confirm that the test fails for the relevant reason on the baseline rather than because setup, imports or fixtures are broken.
  • Confirm separately that it passes on the patch, if both executions are authorised and feasible.
  • Record exact commands and outputs in the next ledger; do not paraphrase a failure into a stronger result.
  • Ask a human to decide whether the assertion accurately represents the public contract. Codex must not self-approve the test or patch.
  • If only the patched state was tested, narrow the claim to “passes on the tested patch state”; do not call it regression proof.
Two separated workspaces containing matching project components, with guarded resource channels and a human approval bridge
The composition reinforces the difference between a maintainer-controlled local workspace and a separately bounded cloud task without implying shared credentials or automatic permission.

Repository-prescribed checks: preserve commands, scope and omissions

Project-defined checks carry more evidential weight than improvised substitutes because they encode repository conventions, but a passing suite still supports only its tested scope. Instructions may be layered through AGENTS.md files from repository root towards the working directory. OpenAI’s current documentation, accessed 2 October 2026, states that combined guidance has a default 32 kibibytes limit. Guidance may therefore be stale, conflicting, incomplete or truncated. Compare it with package scripts, contribution documentation and continuous integration configuration rather than assuming it is exhaustive.

Prompt 14: Run prescribed validation commands

Purpose

Execute only authorised, repository-prescribed validation and produce an exact-command ledger. This preserves the difference between a command that passed, a command that was skipped and a check that does not exist.

Copy-paste prompt

You are validating an authorised SDK parsing patch candidate in a reviewable repository context.

Goal:
Identify and, where authorised, run the repository-prescribed checks relevant to the changed paths. Produce an exact-command ledger without converting partial results into a release recommendation.

Authorised evidence:
- Baseline and patch identifiers: [REFERENCES]
- Changed paths: [PATHS]
- Applicable AGENTS.md files: [PATHS]
- Contribution and test documentation: [PATHS]
- Package scripts or task definitions: [PATHS]
- Existing continuous-integration configuration: [PATHS]
- Approved runtime/tool versions: [VERSIONS OR “NOT SUPPLIED”]

Constraints and stop conditions:
- Do not invent commands, executions, outputs, timings, coverage, compatibility assurances or successful results.
- Do not paste, print or request raw secrets, private vulnerability material, customer data or protected logs.
- Do not access a network, install undeclared tools, change lockfiles, mutate external services or leave the repository boundary without explicit approval.
- Do not use danger-full-access as a default.
- Do not choose deprecation or release timing.
- Do not approve your own patch, merge it or describe checks as proof of correctness, security or release safety.
- Stop before any destructive, publishing, registry, signing or credential-dependent command.

Required checks:
1. Map each changed path to repository-prescribed unit, lint, type, build and packaging checks.
2. Record the command exactly as entered, working directory, approved environment identifier and exit status.
3. Preserve a concise output excerpt or repository-local artefact reference; redact sensitive data rather than placing it in the prompt.
4. Label every relevant check PASS, FAIL, SKIPPED, BLOCKED or NOT FOUND.
5. State why a skipped or blocked check was not run and who must resolve it.
6. Note any environment drift, generated-file change or unexpected working-tree mutation.
7. Recheck git status after execution.

Fixed output artefact:
An exact-command check ledger, followed by omissions, environment caveats, unexpected mutations, evidence-supported conclusions, and questions reserved for an authorised human.

Required inputs

  • Supported source facts: Codex CLI can run local code and supports repeatable automation, while local permissions remain under user control. OpenAI’s sandbox guidance distinguishes filesystem and network boundaries from approval policy; workspace-write with on-request is described as a lower-risk local automation preset.
  • Repository evidence: scripts, task-runner configuration, contribution instructions, continuous integration jobs, runtime declarations and the actual changed paths.
  • Human decisions: permission to install dependencies or use the network, acceptance of environment differences, treatment of flaky checks and whether blocked platform-specific checks must run elsewhere.

Expected output

The ledger should retain literal commands. The following is a worked format example, not a report of product or repository results:

Check Exact command State Evidence Claim boundary
Targeted regression [PACKAGE-MANAGER] test -- [TEST-PATH] --runInBand NOT RUN No execution supplied No behavioural claim available
Unit suite [REPOSITORY-PRESCRIBED UNIT COMMAND] BLOCKED [MISSING APPROVED RUNTIME] Unit-suite status unknown
Lint [REPOSITORY-PRESCRIBED LINT COMMAND] PASS [LOCAL ARTEFACT OR SANITISED OUTPUT REFERENCE] Only the invoked lint configuration passed in the recorded environment
Package build [REPOSITORY-PRESCRIBED BUILD COMMAND] SKIPPED Requires maintainer-approved dependency retrieval No build or distributable claim

Do not replace [REPOSITORY-PRESCRIBED UNIT COMMAND] with a plausible command. Discover it from authorised files. If two instructions disagree, record both and stop for resolution. Running the easier command is not a valid way to resolve a policy conflict.

Verification checkpoint

  • A human compares every exact command with repository instructions and continuous integration configuration.
  • The working tree is inspected for generated, lockfile or snapshot changes; unexplained mutations become findings.
  • PASS means only that the recorded command exited successfully under the recorded conditions. It does not imply complete correctness or support for other environments.
  • Any skipped required check remains a release-governance issue; Codex cannot waive it.
  • For consequential security, privacy, financial, employment, government or other high-impact decisions, require qualified human review and the project’s established process rather than relying on generated analysis.

Dedicated review: assess the patch without self-approval

A dedicated review is a different operation from asking the patch-producing context to summarise its own work. The current Codex CLI page, accessed 2 October 2026, says it can perform a dedicated review without changing the working tree. That is useful for separating review from editing, but it is not independence in the institutional sense and does not create approval authority.

Prompt 15: Review the patch against the public contract

Purpose

Review the diff against documented behaviour, typing, errors and repository rules. Return prioritised findings with evidence paths, while leaving acceptance and disposition to authorised reviewers.

Copy-paste prompt

Perform a dedicated, read-only review of this authorised SDK patch candidate against its public contract.

Goal:
Identify actionable defects, compatibility risks, error-contract changes, public-type changes, documentation mismatches and missing tests. Do not edit the working tree.

Authorised evidence:
- Base reference: [BASE]
- Patch reference or diff: [PATCH]
- Public documentation and examples: [PATHS]
- Export/type declarations: [PATHS]
- Supported-version policy: [PATH]
- Applicable AGENTS.md review rules: [PATHS]
- Check ledger: [PATH OR CONTENT]
- Approved issue summary: [SANITISED SUMMARY]

Constraints and stop conditions:
- Do not invent findings, command results, compatibility assurances, exploit details or repository policies.
- Do not request raw secrets, private vulnerability reports, customer information or proprietary logs.
- Do not change files, fix findings, merge, approve or recommend release merely because listed checks passed.
- Do not select deprecation or release timing.
- Cite file paths, symbols and diff locations for every finding.
- If security-sensitive analysis would require prohibited or unavailable detail, stop and state the authorised defensive evidence needed. Safety checks may delay or withhold a request even when it is framed as defensive and authorised.

Required checks:
1. Compare changed behaviour with documented inputs, outputs and errors.
2. Inspect exported symbols, public types, defaults and serialised forms.
3. Look for tests that assert implementation details rather than public behaviour.
4. Compare the diff with the approved plan and identify scope drift.
5. Check whether examples or documentation now disagree with code.
6. Rank findings by user impact and confidence, not rhetorical severity.
7. Separate defects, questions and optional improvements.

Fixed output artefact:
A prioritised review report with finding ID, priority, confidence, evidence location, affected public contract, consequence, suggested verification, and human disposition field. End with residual unknowns and “approval not granted”.

Required inputs

  • Supported source facts: Codex can perform a dedicated review without changing the working tree. Codex GitHub review rules may be expressed in AGENTS.md (a repository instruction file), but OpenAI’s documentation says those rules do not replace tests, branch protections or required approvals.
  • Repository evidence: base and patch references, public documentation, exported types, approved plan, applicable rules and the command ledger.
  • Human decisions: finding severity policy, whether an ambiguity is a defect, accepted residual risk, reviewer assignment and final approval.

Expected output

A prioritised findings table should make actionability visible. This worked example uses placeholders and does not assert an actual defect:

Priority Example finding Evidence required Disposition owner
In this fictional triage table, P1 is the highest priority, P2 medium and P3 lower; these labels are local to the example, not a universal severity scale.
P1 F-01: proposed parser branch may alter a documented error into a successful value for [INPUT CLASS]. Diff location, documented error contract and focused baseline/patch test Maintainer and required reviewer
P2 F-02: exported type at [PATH:LINE] may not represent the new return shape. Public declaration, build/type-check evidence and consumer compile scenario Public API owner
P3 F-03: example at [DOC PATH] may need clarification after approval. Approved behavioural decision Maintainer or documentation reviewer

Priority should reflect potential user consequence and blocking status. Confidence records evidential strength. A high-impact but low-confidence concern is still worth investigation, but it should not be stated as a confirmed defect.

Verification checkpoint

  • Another authorised human checks each finding against the cited location and rejects unsupported inferences.
  • Review comments are not silently resolved by the same generated review that raised them.
  • Tests, branch protection and required approvals remain operative even when no findings are returned.
  • “No findings” means no findings were identified within the supplied scope; it is not approval or proof of safety.

Compatibility scenarios: sample consumers rather than universal assurances

Compatibility is relational: a package version interacts with consumer code, runtime versions, configuration, data and documented expectations. A patch can satisfy its regression test yet break a wrapper that relied on a public type or error form. Scenario review samples named consumer patterns; it cannot justify “fully backwards compatible” unless the project has a separately defined and adequately evidenced standard for that conclusion.

Prompt 16: Perform a compatibility-scenario review

Purpose

Translate the public contract and support policy into explicit consumer scenarios, then record which combinations were exercised, reasoned about, blocked or left unknown.

Copy-paste prompt

Review this authorised SDK parsing patch through specified consumer scenarios.

Goal:
Build and, where authorised, exercise a bounded compatibility matrix covering documented direct use, public typing, documented errors and declared supported runtimes. Do not claim universal backwards compatibility.

Authorised evidence:
- Patch and base: [REFERENCES]
- Supported-version policy: [PATH]
- Public examples: [PATHS]
- Export/type declarations: [PATHS]
- Approved fixtures: [PATHS]
- Existing compatibility jobs: [PATHS]
- Check ledger and review findings: [REFERENCES]

Constraints and stop conditions:
- Do not invent consumer projects, executions, results, support policies or compatibility assurances.
- Do not use private downstream code, customer logs, raw secrets or unreleased vulnerability details.
- Do not access external repositories or services unless separately authorised.
- Do not treat undocumented accidental behaviour as guaranteed policy; flag it for human interpretation.
- Do not choose semantic versioning, deprecation scope or timing.
- Do not self-approve the patch or infer release safety from sampled scenarios.

Required checks:
1. Include the reported documented parsing case.
2. Include the nearest documented unchanged case.
3. Include documented invalid input and error behaviour where applicable.
4. Include public type/compile consumption if the package exposes types.
5. Cross the scenarios only with repository-declared supported runtimes available in the authorised environment.
6. Record execution state and exact evidence for each cell.
7. List known consumer patterns not represented.

Fixed output artefact:
A consumer-scenario matrix with scenario, contract source, baseline state, patch state, runtimes covered, result status, interpretation, exclusions, and human compatibility decision required.

Required inputs

  • Supported source facts: Codex may inspect and run authorised repository code, but no supplied OpenAI source supports treating sampled executions as a compatibility guarantee.
  • Repository evidence: declared support matrix, examples, public types, compatibility jobs and approved fixtures.
  • Human decisions: which consumer patterns are material, how undocumented reliance is treated, whether support policy permits a patch, and whether additional downstream testing is required.

Expected output

Consumer scenario Contract source Baseline Patch Runtime coverage Bounded interpretation
Documented parser call with [APPROVED FIXTURE] [DOC PATH/SECTION] NOT RUN NOT RUN [DECLARED VERSION PLACEHOLDERS] Required regression scenario; no result yet
Nearest documented unchanged input [TEST OR EXAMPLE PATH] UNKNOWN UNKNOWN None recorded Guards against over-broad parsing change
Documented invalid input [ERROR CONTRACT PATH] UNKNOWN UNKNOWN None recorded Error shape and message stability require review
Consumer compilation against exported type [TYPE DECLARATION PATH] BLOCKED BLOCKED [MISSING TOOLCHAIN] Public typing remains unresolved
Undocumented wrapper relying on prior accidental output No public contract identified NOT TESTED NOT TESTED Not applicable Policy question, not evidence of guaranteed support

Prioritise documented direct use, declared runtimes and known public types. Add a downstream scenario only when its code or fixture is authorised and relevant. Do not copy a private consumer implementation into the repository merely to make the matrix look comprehensive.

Verification checkpoint

  • Compare runtime rows with the repository’s declared support policy, not with guessed “common” versions.
  • Ensure UNKNOWN, BLOCKED and NOT TESTED cells remain visible.
  • Require a maintainer to decide whether the sampled coverage is sufficient for the proposed release class.
  • If a scenario changes, decide whether that is a bug fix, clarified ambiguity or breaking change; the matrix cannot decide policy by itself.

Deprecation design: propose migration without announcing policy

Deprecation is a public commitment, not merely a warning string. It needs a named behaviour, replacement, communication route, support window and removal decision. Codex can organise options from repository policy, but it must not set timing or a semantic versioning outcome unilaterally.

Prompt 17: Design a proposed deprecation path

Purpose

Prepare a reviewable deprecation proposal when preserving the current behaviour indefinitely may conflict with the documented contract or future design. The proposal must leave dates, release numbers and final wording as explicit human decisions.

Copy-paste prompt

Prepare a deprecation proposal for the authorised SDK parsing issue. This is a policy draft, not an announcement or implementation instruction.

Goal:
Describe how consumers could move from [CURRENT BEHAVIOUR] to [PROPOSED BEHAVIOUR] while preserving a reviewable transition and identifying where repository policy is silent.

Authorised evidence:
- Public contract and examples: [PATHS]
- Compatibility matrix: [REFERENCE]
- Supported-version and deprecation policy: [PATH OR “NOT FOUND”]
- Patch options: [REFERENCES]
- Existing warning conventions: [PATHS]
- Contributor attribution preference: [APPROVED PUBLIC INFORMATION]

Constraints and stop conditions:
- Do not invent policy, dates, release numbers, results or compatibility assurances.
- Do not include raw secrets, private reports, personal data or unreleased vulnerability details.
- Do not promise continued support beyond documented policy.
- Do not choose semantic versioning, removal timing, backports or publication timing.
- Do not implement, announce, merge or approve the proposal.
- If the change is security-sensitive, keep public wording non-operational and route disclosure details through the project’s approved human process. Authorised defensive framing does not remove OpenAI safety checks, which may delay or withhold a request.

Required checks:
1. Name the exact behaviour proposed for deprecation.
2. Identify an available migration path or state that none is yet validated.
3. Draft warning text that avoids claiming an approved date.
4. Provide timeline placeholders tied to human-approved milestones.
5. Identify documentation, type, runtime and warning-mechanism consequences.
6. Compare immediate fix, staged deprecation and major-version alternatives.
7. List every decision requiring maintainer or governance approval.

Fixed output artefact:
A deprecation proposal with rationale, affected contract, migration example, warning-text candidate, timeline placeholders, alternatives and trade-offs, compatibility unknowns, approval matrix, and explicit “not approved/not announced” status.

Required inputs

  • Supported source facts: Codex can inspect repository policy and draft reviewable artefacts; the assigned sources do not give it authority to select release class, deprecation policy or timing.
  • Repository evidence: public contract, compatibility matrix, existing policy, warning conventions and validated alternatives.
  • Human decisions: whether deprecation is appropriate, semantic versioning treatment, dates or release milestones, support period, backports, disclosure and final public wording.

Expected output

A useful proposal makes missing authority impossible to overlook:

Behaviour
Example: [EXACT LEGACY PARSING BEHAVIOUR], with the contract location and affected scenarios.
Replacement
Example: use [VALIDATED ALTERNATIVE]. If it has not been tested, label it “proposed; validation pending”.
Warning candidate
“Example wording: [BEHAVIOUR] is proposed for deprecation. Use [ALTERNATIVE]. Removal timing has not been approved.”
Timeline
[DECISION DATE] → [FIRST ELIGIBLE RELEASE] → [MINIMUM SUPPORTED WINDOW] → [EARLIEST HUMAN-APPROVED REMOVAL RELEASE].
Alternatives
Immediate correction preserves the stated contract sooner but may disrupt consumers relying on accidental behaviour; staged deprecation gives migration time but prolongs divergent behaviour; a major-version change creates a clearer boundary but increases release and migration scope.
Status
Not approved, not scheduled and not announced.

Verification checkpoint

  • A maintainer confirms that the replacement is real, documented and testable.
  • Governance owners fill timeline placeholders; generated dates must not enter an issue, warning or release note.
  • Reviewers check warning wording for precision, localisation conventions and noise implications.
  • Security, privacy, legal, financial, employment, government or similarly consequential effects require the appropriate human specialists and project procedures.
  • If no viable migration exists, do not disguise that gap with a timeline; carry it forward as unresolved.

Contributor-facing transparency: report evidence without overclaiming

A contributor reply should explain what was confirmed, what remains uncertain, whose work is being acknowledged and what happens next. Public transparency does not require disclosing private logs, identities, credentials or sensitive vulnerability details. Attribute only information approved for public use.

Prompt 18: Draft an evidence-bounded contributor reply

Purpose

Draft a public issue or pull-request update that reports the current evidence state without implying acceptance, compatibility, release class or timing.

Copy-paste prompt

Draft a contributor-facing update for this authorised SDK parsing issue or pull request.

Goal:
Explain the patch candidate’s current evidence state in plain British English. Separate confirmed evidence, uncertainty, attribution, privacy limits and next action.

Authorised evidence:
- Public issue or pull-request reference: [REFERENCE]
- Sanitised reproduction record: [REFERENCE]
- Regression-test proposal/results: [REFERENCE]
- Exact-command ledger: [REFERENCE]
- Review findings: [REFERENCE]
- Compatibility matrix: [REFERENCE]
- Deprecation proposal, if any: [REFERENCE]
- Approved contributor name/handle and attribution wording: [PUBLIC DETAILS OR “DO NOT NAME”]

Constraints and stop conditions:
- Do not invent results, dates, release numbers, support commitments or compatibility assurances.
- Do not expose raw secrets, private vulnerability information, customer data, private logs, email addresses or other unnecessary personal data.
- Do not blame the reporter or attribute intent.
- Do not promise merge, backport, release, deprecation timing or publication.
- Do not self-approve the patch or present generated review as required human approval.
- If public wording could disclose a security-sensitive technique or uncoordinated vulnerability detail, stop and request the project’s authorised disclosure reviewer. Defensive intent does not remove OpenAI safety safeguards; a request may be delayed or withheld.

Required checks:
1. State confirmed observations and cite public references or repository paths where appropriate.
2. Put every unrun, blocked or uncertain check in a separate uncertainty paragraph.
3. Use only approved public attribution.
4. Explain omitted details through a concise privacy or disclosure boundary without hinting at hidden conclusions.
5. Name the next human-controlled action and owner role.
6. State that release class and timing remain undecided unless an authorised published decision says otherwise.
7. Avoid “fixed”, “safe” and “backwards compatible” unless the supplied evidence and authorised human wording support the exact claim.

Fixed output artefact:
A contributor reply with labelled sections: Confirmed evidence; Uncertainty; Attribution; Privacy and disclosure; Next action. Follow it with a private maintainer note listing unsupported claims removed from the public draft and decisions still required.

Required inputs

  • Supported source facts: OpenAI’s current guidance supports reviewable prompting and review, but generated text is not maintainer approval. Additional safety checks may delay or withhold cybersecurity-sensitive requests.
  • Repository evidence: sanitised public records, test and command ledgers, review findings, compatibility scenarios and approved attribution details.
  • Human decisions: what may be disclosed, who may be named, whether the patch is accepted, next owner, release class, timing and any coordinated security communication.

Expected output

Confirmed evidence: Example only: “We have a patch candidate for the documented parsing case at [PUBLIC REFERENCE]. The focused check [EXACT COMMAND] has status [PASS/FAIL/NOT RUN] in [RECORDED ENVIRONMENT]. This evidence is limited to the listed fixture and environment.”

Uncertainty: “The [RUNTIME OR CONSUMER] scenario remains [BLOCKED/NOT TESTED], and [REVIEW FINDING] still requires disposition. We are not yet asserting general backwards compatibility.”

Attribution: “Thank you to [APPROVED PUBLIC HANDLE] for reporting the documented case.” If naming is not approved: “Thank you to the contributor who supplied the public reproduction.”

Privacy and disclosure: “We have omitted non-public logs and environment details that are unnecessary to evaluate the public behaviour.” Do not imply that omitted data confirms the fix.

Next action: “An authorised maintainer will review [OPEN FINDING OR CHECK]. Merge, release class and timing remain undecided.”

The private maintainer note should list edits such as: removed “fully fixed” because only one fixture was exercised; removed a proposed release date because governance has not approved it; withheld a private log excerpt because the public claim does not require it.

Verification checkpoint

  • The maintainer traces every public factual sentence to a public or safely sanitised record.
  • The named contributor and attribution wording are approved; otherwise use a neutral description.
  • Uncertainty is visible in the public reply rather than hidden only in the private note.
  • A human disclosure owner reviews security-sensitive wording before publication.
  • The reply does not imply merge, release or deprecation approval.

Carry unresolved evidence into release governance

At this stage, the maintainer has six possible artefacts: causal regression coverage, an exact-command ledger, dedicated review findings, a compatibility-scenario matrix, a deprecation proposal and a contributor reply. None is a release decision. Codex review and AGENTS.md rules do not replace tests, branch protection or required approvals. A passing command does not erase a blocked scenario; a clean review does not waive repository policy; a proposed warning does not establish deprecation timing.

Transfer unresolved items without softening them. At minimum, carry forward: any baseline comparison not run; prescribed checks marked FAIL, SKIPPED, BLOCKED or NOT FOUND; environment drift; unexplained working-tree changes; unresolved review findings; unsupported runtimes or consumer scenarios; ambiguity in the public contract; an unvalidated migration path; missing deprecation policy; attribution or disclosure approval; and every undecided merge, backport, semantic versioning, release-timing, tagging, publishing or security-disclosure question.

Any item that could materially change the public contract, compatibility assessment, disclosure posture or release class remains open until an authorised human records its disposition and supporting evidence. The next section may organise those issues into release governance, but it must not convert absence of evidence into approval.

Validated maintenance evidence can now be converted into public-facing candidates, but not into a release instruction. In the fictional SDK parsing-regression example followed through this playbook, Codex may organise repository evidence, draft wording and expose unresolved choices. It must not choose semantic versioning, merge the pull request, finalise the changelog, create a tag, publish to a registry, announce a vulnerability or decide who receives acknowledgement. Those actions require an authorised maintainer who understands the project’s policy, downstream impact and disclosure obligations.

The prompts below follow OpenAI’s documented goal, context, constraints and “done when” pattern, with a fixed output artefact added for this playbook. Every citation must point to a repository path, commit, diff, command record or supplied artefact. A generated citation is not evidence until a human opens and checks it. Remove personal data, credentials, private reports and unreleased vulnerability details before supplying inputs. Security, privacy, financial, employment, government and other consequential decisions always require qualified human review.

Prompt 19: Draft evidence-based release-note candidates

Purpose

Translate the validated behavioural change into candidate release-note language without allowing Codex to select a release class. The distinction is between describing what changed and deciding whether that change belongs in a patch, minor or major release. In the parsing-regression example, the evidence might establish only that a particular documented input shape is handled differently; it does not establish universal compatibility.

Copy-paste prompt

You are preparing candidate release-note text for an authorised open-source maintainer. Do not choose a release class, edit the changelog, tag, merge or publish.

Goal:
Draft separate Patch, Minor and Major release-note alternatives for the validated SDK parsing-regression change.

Authorised evidence:
- issue/reproduction artefact: [path or supplied reference]
- approved diff or commit: [reference]
- regression test: [path and test name]
- public contract evidence: [documentation/API paths]
- compatibility scenario record: [path]
- contributor-reply draft: [path, if approved for use]

Constraints and stop conditions:
- Use only claims directly supported by those artefacts.
- Cite a repository path or supplied artefact after every factual claim.
- State unknown affected versions, platforms and consumer patterns.
- Do not infer security impact, backwards compatibility or semantic versioning.
- Do not include private reporter data, raw logs, secrets or undisclosed vulnerability details.
- Stop and list missing evidence if the user-visible effect cannot be established.
- Require authorised human approval for wording, release class, attribution and disclosure.

Required checks:
1. Separate confirmed behaviour from inferred impact.
2. Identify migration action only when a documented action exists.
3. Avoid “fully fixed”, “safe”, “all users” and equivalent universal claims.
4. Explain why each alternative might fit, and what evidence would rule it out.
5. Preserve project terminology from cited public documentation.

Output artefact:
A release-note candidate packet containing:
- one-sentence evidence scope;
- Patch candidate;
- Minor candidate;
- Major candidate;
- uncertainty and omitted-claim list;
- citation list;
- explicit HUMAN DECISION fields for release class, final wording,
  acknowledgement and disclosure.

Required inputs

  • The reviewed patch or commit reference, not an uncommitted recollection.
  • The issue and reproduction records with private data redacted.
  • The exact regression-test path and the behaviour it isolates.
  • Public API, README, reference or compatibility-policy locations.
  • Known validation omissions and the authorised project terminology.

Expected output

A useful packet presents alternatives rather than disguising a recommendation as fact. The following is a worked structure, not a semantic versioning decision or promised product output:

Candidate Example public wording Evidence needed Trade-off and human decision
Patch “Correct parsing of the documented [input form] when [condition] applies. See [test path] and [documentation path].” Evidence that existing documented behaviour regressed and that the patch restores it without intentionally expanding the contract. Least disruptive classification, but unsuitable if the change adds supported behaviour or breaks consumers. A maintainer decides.
Minor “Add or clarify support for [input form], with regression coverage at [test path].” Evidence that the behaviour is additive under the project’s versioning policy, plus documentation of the newly supported contract. Signals a capability addition, but may misrepresent a bug fix. A maintainer decides whether the project’s policy treats it as additive.
Major “Change parsing of [input form]; consumers relying on [old behaviour] should follow [migration path].” Confirmed incompatible behaviour, identified affected contract and a reviewed migration route. Makes disruption explicit but must not be chosen merely because impact is uncertain. A maintainer approves timing, migration and deprecation policy.

Verification checkpoint

  • Supported source facts: every behavioural sentence resolves to an opened path, diff, test or supplied artefact; omissions are visible.
  • Human decisions: release class, changelog wording, attribution, timing and disclosure remain blank or explicitly marked for approval.
  • Reject the packet if it converts a passing test into a broad compatibility claim or uses absence of failures as proof of safety.

Prompt 20: Prepare documentation deltas

Purpose

Locate documentation that may need to change while distinguishing a documentation candidate from an authorised publication. This catches a common mismatch: code and tests describe one contract while README examples or generated references describe another.

Copy-paste prompt

Goal:
Prepare a documentation-delta checklist for the approved parsing patch.
Do not edit, publish or regenerate documentation unless separately authorised.

Context:
- approved diff: [reference]
- contract-impact map: [path]
- release-note candidates: [path]
- repository documentation instructions: [AGENTS.md paths]
- docs build/check commands: [repository source]

Constraints and stop conditions:
- Cite exact repository paths and, where possible, headings or symbols.
- Label generated files and identify their source; do not hand-edit generated
  output unless repository instructions require it.
- Do not invent examples, support promises, migration dates or release numbers.
- Redact private information and exclude secrets.
- Stop when the authoritative documentation source cannot be identified.
- Require human approval before edits, publication or policy wording.

Required checks:
- README quick-start and examples;
- API/reference source;
- exported types, docstrings or command help;
- migration/deprecation guidance;
- version-support pages;
- changelog source versus generated output;
- translation or duplicated examples;
- documentation validation command and any skipped checks.

Output artefact:
A path-level checklist with current claim, proposed delta, supporting evidence,
owner/approval field, generated-source warning, validation step and uncertainty.

Required inputs

Supply the contract map, approved patch and applicable layered AGENTS.md guidance. OpenAI’s current guidance, accessed 2 October 2026, says Codex reads such files from the repository root towards the working directory and applies a default 32 kibibytes combined-guidance limit. Consequently, verify the effective instructions manually: they may be stale, conflicting, incomplete or truncated.

Expected output

Location check Question Required evidence Decision rule
README or quick-start Does a copied example exercise the changed parser path? Exact heading and example path. Propose a delta only if the example’s interpretation changes.
API reference source Does the parameter, return value or error contract differ? Source symbol and generated-reference mapping. Edit the source of truth, not generated output, unless instructions say otherwise.
Examples and fixtures Would an example teach the old behaviour? Example path plus test or runnable check. Prefer an existing verified example; label new sample output as illustrative.
Migration guide Must consumers alter code or data? Confirmed incompatible scenario. No migration claim for a restorative patch unless evidence shows consumer action.
Deprecation page Is a warning or transition already approved? Maintainer decision record. Leave dates and removal versions as placeholders until approved.
Changelog source Where is release text authored? Repository convention. Prepare a candidate only; finalisation belongs to the release maintainer.
Translations or duplicates Are equivalent claims maintained elsewhere? Path search and ownership note. Flag rather than silently rewrite content outside the approved scope.
Validation What checks links, examples or generated references? Repository-prescribed command. Record exact results and omissions; a pass is not proof of accuracy.

Verification checkpoint

  • Supported source facts: paths exist, quoted claims match their source and generated files have an identified origin.
  • Human decisions: publication, migration policy, deprecation dates, version labels and documentation ownership are explicitly reserved.
  • If code behaviour is validated but public wording remains disputed, retain the documentation delta as unresolved rather than making the code determine policy.

Prompt 21: Build the release-candidate evidence ledger

Purpose

Consolidate traceable evidence without laundering uncertainty into a “ready” verdict. A ledger records what was observed, where it can be checked and who must decide. It is not a safety certificate.

Copy-paste prompt

Goal:
Build a release-candidate evidence ledger for maintainer review.

Inputs:
[issue artefact], [reproduction record], [contract map], [approved plan],
[diff/commit], [test ledger], [compatibility scenarios], [docs checklist],
[release-note candidates], [review findings].

Constraints:
- Do not merge, tag, publish, backport or mark the candidate approved.
- Cite repository paths, commit identifiers or supplied artefacts.
- Separate observation, interpretation and human decision.
- Preserve failed, skipped and unavailable checks.
- Do not include secrets, personal data or private vulnerability content.
- Stop if evidence refers to a different revision or cannot be opened.
- Require explicit human approval for consequential decisions.

Required checks:
Confirm revision identity, evidence provenance, command scope, unsupported
environments, documentation status, residual compatibility risk, reviewer
independence and unresolved disclosure questions.

Output artefact:
A ledger using the stated schema, followed by blockers, non-blocking unknowns,
and numbered approval fields. Never output “safe to release”.

Required inputs

Use immutable or clearly identified revisions where available, exact command records and redacted issue material. Do not paste registry tokens, network credentials, private customer logs or embargoed vulnerability details into a prompt.

Expected output

Ledger field Content Acceptance rule
Candidate identity Base revision, candidate revision, branch/worktree and dirty-state note. All later evidence must refer to this candidate or explain divergence.
Issue scope Redacted issue reference, classification and authorised boundary. No private or unrelated material.
Reproduction Environment, fixture, commands and observed result. Distinguish reproduced, not reproduced and not attempted.
Public contract Relevant documented API, examples, types or flags. Path citations opened by a reviewer.
Change Files, semantic diff and deviations from the approved plan. Unapproved scope expansion is a blocker.
Checks Exact commands, results, failures, skips and limitations. No collapsed “all checks passed” summary without command detail.
Compatibility Scenarios actually examined and populations not represented. No universal backwards-compatibility claim.
Documentation Required deltas, owners and unresolved wording. Publication state stated accurately.
Review Reviewer basis, findings and dispositions. Self-review and independent review identified separately.
Residual risk Known unknowns and consequence if wrong. Owner and decision point assigned.
Human approvals Patch acceptance, release class, merge, tag, publish, backport, acknowledgement and disclosure. No approval inferred from silence or tool output.

Verification checkpoint

  • Supported source facts: revision identity, paths, commands and observed results are reproducible from the cited records.
  • Human decisions: every release action has a named approval field and remains pending unless an authorised person records it.
  • Fail the checkpoint if skipped checks disappear, evidence mixes revisions, or a generated summary is the sole source for a claim.

Prompt 22: Request an independent local review

Purpose

Ask for a fresh review against the base revision without changing the working tree. OpenAI’s Codex CLI page, accessed 2 October 2026, describes dedicated review without modifying the working tree. That capability does not replace tests, branch protection or required approvals.

Copy-paste prompt

Goal:
Perform an independent review of [candidate revision] against [base revision]
without editing files.

Context:
Apply the verified repository instructions at [paths]. Review the approved plan,
contract map, tests and evidence ledger, but do not accept their conclusions
without checking the cited sources.

Constraints:
- Read-only review; no fixes, commits, merges or publishing.
- Stay within the authorised repository and local tool permissions.
- Cite file path and line/range or artefact for every finding.
- Distinguish confirmed defect, plausible concern and question.
- Do not expose secrets or private vulnerability information.
- Stop if the base/candidate identity is ambiguous.
- Require human review of security, privacy and release implications.

Required checks:
Scope drift, parser behaviour, error handling, public types, compatibility,
test causality, documentation mismatch, omitted checks and release-note accuracy.

Output artefact:
Prioritised findings plus disposition candidates. Include “no finding” only for
areas actually inspected; do not provide a release approval.

Required inputs

Provide base and candidate identifiers, effective instructions, the approved scope and commands the reviewer may run. Local tools and permissions remain user-controlled. Keep sandbox boundaries distinct from approval policy: OpenAI’s current sandbox guidance describes workspace-write with on-request as a lower-risk local automation preset, but project and managed-workspace rules vary. Do not default to danger-full-access.

Expected output

The finding log should separate severity from disposition. Severity describes possible consequence; disposition records what a human chooses after checking it.

Finding status Meaning Required disposition evidence
Confirmed The cited diff and contract establish a concrete mismatch. Fix reference, scope change approval or explicit human decision not to proceed.
Plausible concern A credible risk exists, but evidence is incomplete. Targeted check, documented limitation or authorised risk decision.
Question Policy, intent or environment cannot be inferred. Maintainer answer linked to the relevant artefact.
False positive The cited evidence does not support the concern. Path or test demonstrating why; not a bare dismissal.
Duplicate Another finding covers the same cause. Reference to the retained finding.
Deferred Valid but outside this candidate’s approved scope. Tracked follow-up and human acceptance of deferral.

Verification checkpoint

  • Supported source facts: each finding names inspected evidence and does not claim coverage beyond inspected files and checks.
  • Human decisions: severity acceptance, deferral, scope expansion and release blocking remain with maintainers.
  • A clean review means only that no finding was recorded within its stated scope. It does not prove correctness, security or release safety.
  • Record uncertainty and inconclusive review findings in the release evidence ledger; a blank finding list is not a proof that hidden regressions or security faults are absent.

Prompt 23: Prepare numbered maintainer release questions

Purpose

Turn unresolved trade-offs into answerable human questions. This avoids the failure mode in which a long report hides a release decision inside apparently factual prose.

Copy-paste prompt

Goal:
Convert unresolved ledger items into numbered questions for authorised maintainers.

Context:
Use [evidence ledger], [review findings], [release-note packet] and [docs checklist].

Constraints:
- Do not answer policy questions or recommend publication as a conclusion.
- For each question, cite supporting evidence and state consequences of each option.
- Separate missing facts from value or policy judgements.
- Do not include secrets, identities or private disclosure details.
- Escalate security, privacy, financial, employment, government or other
  consequential matters to qualified human reviewers.
- Require explicit recorded approval; silence is not consent.

Required checks:
Ask about semantic versioning, backport, deprecation, timing, documentation,
acknowledgements, disclosure, residual failures and rollback/revert readiness.

Output artefact:
A numbered decision sheet with options, evidence, uncertainty, decision owner,
deadline placeholder and “no decision recorded” default.

Required inputs

Supply only the redacted ledger and unresolved findings. The owner list must come from project governance rather than being invented by Codex.

Expected output

  1. Which release class, if any, matches the project’s documented policy for this change?
  2. Must the patch be backported, and which maintained branches are eligible?
  3. Does any changed behaviour require deprecation rather than immediate replacement?
  4. Are documentation deltas required before merge, before publication, or later?
  5. Do failed, skipped or unavailable checks block this candidate?
  6. Is contributor acknowledgement appropriate, and has the proposed wording protected privacy?
  7. Does the issue require a private disclosure process or coordinated communication?
  8. Who may merge, finalise the changelog, create the tag and publish to each registry?
  9. What evidence would trigger a revert, withdrawal of a published release or a follow-up release?

Verification checkpoint

  • Supported source facts: every question is tied to a real unresolved ledger entry.
  • Human decisions: answers, owners, deadlines and risk acceptance are recorded by authorised people.
  • Remove questions already settled by cited policy; retain questions that require judgement rather than allowing Codex to infer intent.

Prompt 24: Draft the pull request and release handoff

Purpose

Prepare a reviewable pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry body and release handoff while preserving operational gates. A handoff packages evidence; it does not execute release operations.

Copy-paste prompt

Goal:
Draft a PR body and release handoff for the parsing-regression candidate.

Context:
Use only [ledger], [review dispositions], [docs checklist], [release-note
candidates] and recorded maintainer answers.

Constraints:
- Cite paths, commits and artefacts.
- Do not claim approval that is not recorded.
- Do not merge, tag, finalise the changelog, publish, backport, acknowledge or disclose.
- Never include tokens, private logs, reporter identities or embargoed details.
- Mark unknowns and failed/skipped checks prominently.
- Require authorised human approval at every operational gate.

Required checks:
Revision identity, issue scope, behaviour before/after, contract impact, tests,
validation omissions, compatibility limits, docs state, review findings,
rollback/revert evidence and outstanding decisions.

Output artefact:
1. Draft PR title and body.
2. Candidate release-note references.
3. Human-controlled handoff checklist.
4. Blockers and residual unknowns.
5. Explicit statement that Codex has not approved or executed a release.

Required inputs

Use the same candidate revision throughout. If a new commit follows review, regenerate or amend the evidence instead of attaching stale results.

Expected output

  • PR scope and non-goals cite the approved plan.
  • Before/after behaviour cites the reproduction and regression test.
  • Validation lists exact commands, failures, skips and environment limits.
  • Compatibility claims are scenario-bounded.
  • Documentation work is marked complete, pending or not applicable with evidence.
  • Review findings link to dispositions; deferred findings link to follow-up tracking.
  • Authorised human: approve the patch.
  • Authorised human: choose the release class.
  • Authorised human: merge under branch-protection and review rules.
  • Authorised human: finalise the changelog and release notes.
  • Authorised human: create and sign the tag where project policy requires it.
  • Authorised human: publish to each registry using approved credentials and procedures.
  • Authorised human: decide and execute any backport.
  • Authorised human: approve acknowledgements and protect contributor privacy.
  • Authorised human: decide any security disclosure and communications route.

Verification checkpoint

  • Supported source facts: the PR draft matches the candidate revision and preserves adverse evidence.
  • Human decisions: all operational boxes remain unchecked until an authorised person acts.
  • Do not use a generated handoff to bypass existing continuous integration, branch protections, required reviews or registry controls.

Prompt 25: Propose one durable post-release instruction improvement

Purpose

Capture one verified recurring lesson after human resolution, without automatically rewriting repository policy. Durable guidance belongs in AGENTS.md only when it is stable, concise and relevant to future work; issue-specific facts generally belong in tests, documentation or maintainer records.

Copy-paste prompt

Goal:
Propose exactly one durable instruction or maintainer-checklist improvement
derived from the resolved parsing-regression work.

Context:
Use the final human decision record, accepted patch, reviewed evidence and
post-release observations supplied by the maintainer.

Constraints:
- Do not edit AGENTS.md or policy files.
- Do not generalise from an unresolved or one-off event.
- Cite the repository evidence supporting recurrence.
- Check for duplicate or conflicting guidance from root to current directory.
- Keep the proposal concise because combined AGENTS.md guidance has a default
  32 kibibytes limit and may be truncated.
- Exclude secrets, private reports and vulnerability details.
- Require human approval for any policy or instruction change.

Required checks:
Is the lesson durable? Is AGENTS.md the right location? Could a test,
documentation correction or checklist be more enforceable? How will maintainers
verify that the instruction is useful and non-conflicting?

Output artefact:
One proposed change with target path, exact candidate text, rationale,
evidence citations, conflict check, verification method and HUMAN APPROVAL field.

Required inputs

Wait for the authorised resolution. A release candidate alone is insufficient: the premise may change during review, backporting or publication. Post-release observations must be supplied as artefacts rather than invented.

Expected output

Example candidate, not a product guarantee: “When changing parser acceptance rules, update or cite the contract fixture at [path] and record at least one rejected-input case.” Choose a checklist instead if the rule applies only during releases; choose a regression test if machine-enforceable behaviour is the real requirement.

Verification checkpoint

  • Supported source facts: the proposal is traceable to accepted repository evidence and does not duplicate nearer guidance.
  • Human decisions: maintainers decide whether the lesson is durable, where it belongs and whether to adopt it.
  • Reject vague instructions such as “ensure compatibility”. A durable rule must name an action and a reviewable artefact.

Selective routes and failure controls

Issue route Use Skip or adapt Human gate
Restorative patch Prompts 19–25, with patch wording tied to the existing contract. Do not invent migration or deprecation material if consumer action is unnecessary. Confirm that project policy treats the behaviour as a patch.
Potentially breaking change Keep all compatibility, deprecation, documentation, ledger and decision steps. Do not use a patch candidate merely because the code diff is small. Approve contract change, release class, migration, timing and disclosure.
Documentation-only issue Use Prompts 20, 21, 23–25; retain evidence showing code already matches intended behaviour. Omit patch-specific testing only with a recorded reason; run relevant docs checks. Approve authoritative wording and publication.

Stop if reproduction is unverified, the candidate revision changes, evidence comes from unauthorised data, a private security report appears, or requested commands exceed repository and sandbox boundaries. Re-route a support question instead of manufacturing a code change. Treat scope creep, uncited compatibility language, omitted failures, hand-edited generated files and inferred maintainer consent as blockers. A test generated after seeing the patch can still be useful, but its causal claim requires human inspection and, where practical, confirmation that it distinguishes the pre-change and post-change behaviour.

Concise product-boundary questions

Can Codex CLI prepare these artefacts?
OpenAI’s current CLI documentation, accessed 2 October 2026, says it can inspect, edit and run local code, support repeatable automation and perform dedicated review without changing the working tree. Local tools, permissions and release authority remain user-controlled.
Does Codex Cloud share one mutable task workspace?
No. OpenAI’s current Cloud environment documentation, accessed 2 October 2026, says published environments reuse prepared filesystem state and each new task receives its own workspace. Repository access, credentials, network policy and service authorisation remain separate.
Should network secrets be pasted into prompts?
No. The Cloud documentation describes substitution of configured network secrets for allowed HTTPS destinations. Enabling a destination does not itself grant service permission. Keep raw secrets and untrusted private data out of prompts.
Are Chat, Work, Codex and the API interchangeable?
No. OpenAI’s support article updated 2 October 2026 distinguishes Chat, Work and Codex, with Codex intended for software-development and technical work. A chat conversation does not thereby gain local terminal or repository access, and product access does not imply API parity.
Is Cloud available to every reader?
Do not assume so. Plan, role, client, workspace settings and rollout can vary. Check the current account and workspace documentation rather than inferring availability from a subscription name.
Does sandboxing remove the need for approvals?
No. Sandbox mode governs technical file and network boundaries; approval policy governs when Codex pauses for permission. Use least privilege and do not prescribe unrestricted filesystem and network access by default.
Does a green status page guarantee a release task will work?
No. The OpenAI status page’s historical aggregate figures are context, not a service-level agreement or prediction. The OpenAI status page showed 99.95% aggregate Codex uptime for July–October 2026 when reviewed on 2 October 2026. That historical aggregate does not establish individual availability, which can vary by tier, model and feature.
Can defensive security work still encounter safety checks?
Yes. OpenAI states that additional automated safety checks may delay or withhold some biological and cybersecurity requests across ChatGPT, Codex and the API. Authorised or defensive framing does not guarantee a response. Do not place private vulnerability material in prompts; use the project’s approved disclosure process and human security review.

Conclusion

The appropriate end state is not “Codex says release”. It is a revision-specific packet in which public wording, documentation deltas, tests, command results, review findings, unknowns and human decisions remain distinguishable. Codex can assist with inspection, drafting and structured review, but generated evidence and passing checks cannot guarantee correctness, compatibility, security or release safety. Authorised maintainers must inspect the cited sources and retain control of merge, semantic versioning, changelog finalisation, tagging, publication, backports, acknowledgements and disclosure.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this