Codex Prompting Playbook for One Fictional Settings-Persistence Defect: Scoped Patch, Regression, and Human Review

Conceptual illustration of a settings dial appearing to click into place then returning to its former position after a circular refresh path, with no interface, text or logo.

Establish the fictional case before asking Codex to investigate

This settings-persistence defect is fictional. Nothing in this playbook reports a defect in an OpenAI product, the reader’s application or any other deployed service. The workflow is unrun: its route, controls, files, requests, storage mechanism, commands and results remain illustrative placeholders until an authorised engineer replaces them with repository-specific evidence.

The single working example is an authorised development copy of a web application. A test actor visits /settings/notifications, changes Enable alerts, selects Save, sees a success signal and then refreshes the page. The control appears to return to its starting value rather than retaining the requested value. This description does not establish whether the error is in the interface, submission path, validation, write path, reload path, fixture or observation method. Those are competing possibilities to investigate, not facts to insert into a prompt.

OpenAI’s Codex prompting guidance, accessed on 3 October 2026, recommends naming the desired behaviour, relevant code or reproduction, constraints and verification method. The same guide uses an illustrative settings-screen example in which a “Saved” message appears but a changed toggle resets after refresh. OpenAI presents that as an example prompt, not as a documented real-world defect or a guarantee that Codex will find and repair one. This playbook therefore turns the resemblance into a disciplined evidence procedure rather than treating the example as a known diagnosis.

Repeat the original user interface reproduction after the candidate patch. That later repeat is mandatory because source inspection or a test result cannot establish what the original user path now does. It is nevertheless only one evidence type: a successful manual repeat would not, by itself, prove correctness for other users, settings, environments or concurrency conditions.

Do not commit, push, open a pull request, or merge. That sentence belongs in every packet that could otherwise be interpreted as permission to take a repository action. It is an instruction, not an enforceable security boundary. Apply the distinct sandbox, approval, account, data and repository controls set out in Keep product controls distinct from repository governance.

The practical objective for this first phase is narrower than “fix the bug”. Freeze a defect contract, confirm which observations are genuinely available, record omissions and map a safe read-only investigation. Do not ask for edits, a root cause or a patch yet. If the reproduction cannot be stated without assumptions, remain in contract-building mode; if it can be stated but has not been observed in the authorised environment, run or arrange observation before diagnosis; only if the reproduction and context are recorded should repository inspection begin.

Conceptual illustration of a settings dial appearing to click into place then returning to its former position after a circular refresh path, with no interface, text or logo.
A visible save signal is not evidence that a setting survived the reload path.

Evidence checkpoints

Documented point: OpenAI’s prompting guidance says a useful Codex prompt should name desired behaviour, relevant code or reproduction, constraints and a verification method. This does not guarantee root-cause identification or a correct patch. [OpenAI Codex prompting guidance]

Documented point: OpenAI’s own illustrative bug-fix prompt uses a settings-screen case where Saved appears but a changed toggle does not persist after refresh. It is a generic example prompt, not evidence of a real OpenAI or reader-product defect. [OpenAI Codex prompting guidance]

Documented point: The same guidance asks Codex to repeat reproduction steps after a fix and, where relevant, report commands and results for lint and the smallest relevant test suite. Actual command state can be pass, fail, blocked or skipped; no result should be invented. [OpenAI Codex prompting guidance]

Documented point: OpenAI says sandboxing and approvals are separate controls, and identifies the workspace-write sandbox setting with on-request approvals as a lower-risk local automation posture. The chosen local setting is not a universal guarantee; repository and workspace policies may impose additional limits. [OpenAI Codex sandboxing guidance]

Documented point: Codex Cloud eligibility and usage are plan/workspace/rollout dependent, and the opened plan page says Cloud is not included with Free or Go. This article is intentionally local and does not require or promise Cloud. [Using Codex with your ChatGPT plan]

Documented point: The data-controls setting on personal plans also applies to Codex tasks, but an opt-out is not a deletion or repository-security control. Managed workspaces have their own retention/access policies. [Data controls in ChatGPT]

Documented point: The 6 July 2026 release note describes GPT-5.5 Instant Mini as a ChatGPT fallback and says the update does not affect application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry or Codex. The ChatGPT fallback notice does not establish any Codex model selection or API availability. [ChatGPT release notes]

Separate the user-visible symptom from the engineering claim

A user-visible symptom is a bounded observation: after a particular action and refresh, a control displays a particular value. An engineering claim goes further: for example, that a request omitted a field, a server rejected a write, a cache served stale data or a database retained the old value. The first can be entered into the contract when directly observed; the second remains a hypothesis until evidence discriminates it from alternatives.

Use a three-column note before writing any Codex prompt. In the first column, put observations that a named person or approved tool can repeat. In the second, put interpretations that still require evidence. In the third, put unknowns. For the fictional example, “the success signal appeared after Save” may be an observation. “The server accepted the new preference” is an interpretation. “Whether the application reads this preference from a server, local store or fixture” is an unknown.

A practical verification step is to challenge every sentence with: “What authorised artefact would show this?” A screen observation can support what was rendered. A permitted request trace may support what was sent and returned. Source references may support how a code path is intended to work. A controlled read through the application’s normal boundary may support what was durable. No single artefact should silently stand in for all four.

For example, do not write, “Saving the toggle succeeds, but the database is not updated.” Write, “Illustrative observation: the interface displayed the configured success signal after Save; after the specified refresh, the control displayed the starting value. Storage technology and write outcome are unknown.” The second wording preserves what must be explained without inventing an implementation.

A premature implementation story can make an investigation feel efficient, but it steers file searches, tests and proposed edits towards one unproven cause. A symptom-only contract takes slightly longer to prepare but leaves room for a client-only state error, omitted submission field, validation rejection, stale reload, incorrect read mapping or environment problem. Prefer the latter whenever more than one explanation remains compatible with the observations.

Freeze the defect contract with evidence placeholders

The defect contract is a compact, reviewable description of the behaviour to reproduce. It differs from a bug summary because it carries enough environmental and procedural detail for another authorised person to attempt the same observation. It also differs from a diagnosis because it does not explain why the behaviour occurs.

Create the contract in a repository-approved scratch note or task message that contains no secrets, private customer data or production identifiers. Use test data and an authorised non-production actor. Replace each placeholder below; if a value is unavailable, write UNKNOWN and name the person or observation needed to resolve it. Never fill a gap with a plausible framework, endpoint or storage layer.

Route and authorised test actor

Route identifies the exact navigational target and any safe prerequisites needed to reach it. The fictional placeholder is [ROUTE: /settings/notifications in the authorised development copy]. Record whether this is a path, local Uniform Resource Locator (URL)The address used to identify and access a resource on the web. Open glossary entry or application navigation description according to repository practice, but omit live hostnames or query values that expose sensitive data.

Authorised test actor identifies the non-production role or fixture permitted to alter the setting. An acceptable example is [AUTHORISED TEST ACTOR: local fixture with permission to update its own notification preferences]. Do not paste credentials, tokens, session cookies, personal email addresses or copied environment files. If the actor’s authorisation is unclear, stop rather than asking Codex to discover or bypass access.

A human must confirm both that the route belongs to the intended development environment and that the actor may perform the mutation. No authorisation evidence means no save attempt. Read-only source inspection may still be possible under repository policy, but it cannot be used to justify an unapproved data mutation.

Starting persisted value, action and expected transition

Starting persisted value is the value established before the reproduction through an authorised durable read or a controlled fixture setup. It is not merely the toggle’s initial appearance. The example placeholder is [STARTING PERSISTED VALUE: Enable alerts = on, established by the approved fixture/read method]. If only the rendered state is known, label it starting displayed value and keep persisted state unknown.

user interface (UI)The controls and visual surfaces through which a person interacts with software. Open glossary entry action records an exact interaction sequence. An illustrative entry is [UI ACTION: change Enable alerts from on to off, then select Save once]. Include relevant prerequisites such as waiting for initial loading to finish, but do not invent button names, keyboard shortcuts or timing thresholds. If multiple actions can save the form, select one path for this defect contract and list the others as out of scope.

Expected value states the durable behaviour required after the authorised save. In the fictional example it is [EXPECTED VALUE: Enable alerts remains off after the defined refresh/reopen method]. This expected value should come from an accepted requirement, existing behaviour, product decision or test convention—not from Codex guessing what the control ought to do.

Compare three moments separately: the established starting value, the value requested by the action and the value obtained after the defined reload boundary. A defect contract is not complete if “before” refers only to appearance while “after” claims durable storage. Keep both at the same evidential level or state the mismatch explicitly.

Success signal and why it is not a durability assertion

Success signal records exactly what the interface showed after submission. It could be a message, status region or another repository-defined indication. The fictional placeholder is [SUCCESS SIGNAL: the interface displays its configured post-save success indication]. “Saved” may be used only if that is the observed text; do not quote unobserved interface copy.

A toast or comparable message is evidence that the interface rendered a signal. It does not, without additional evidence, prove that a request contained the intended value, that validation accepted it, that a write completed, that a later read returns it or that the refreshed interface maps it correctly. The signal could be tied to optimistic client state, a generic completion branch or a response whose semantics have not yet been inspected.

For example, suppose the control changes from on to off and a success signal appears. The defensible record is: “The interface displayed the success signal after the action.” The indefensible shortcut is: “The new value was persisted and then lost.” The latter invents two durable-state events that the display alone cannot establish.

Verify the distinction by asking Codex, in read-only mode, to identify where the success signal is triggered and to cite repository paths and lines if available. That source observation may show which condition causes the signal, but it still does not establish runtime behaviour until corroborated. If the signal is emitted only after an apparent success branch, record that branch condition precisely; do not paraphrase it as a confirmed storage write.

Treat a success message as a user-interface signal unless an authorised observation crosses the relevant persistence boundary. Even then, retain separate evidence for signal, write and reload. That extra bookkeeping avoids a false conclusion based on optimistic presentation.

Refresh or reopen method, observed value and environment

Refresh/reopen method defines how transient state is discarded and the value is obtained again. The placeholder might be [REFRESH/REOPEN METHOD: perform the agreed browser refresh after the save signal, then wait for the settings view to finish loading]. Depending on the actual application, the correct boundary could instead be route navigation, remounting, a new session or another documented action. Do not combine these as though they were equivalent.

Observed value is what appears after that exact boundary. The fictional example is [OBSERVED VALUE: Enable alerts displays on after refresh]. Record uncertainty if loading, stale rendering or a failed request makes the value ambiguous. A screenshot may support rendered appearance where permitted, but it should not contain personal data and should not be presented as proof of server state.

Browser/runtime identifies the actual authorised execution context, such as [BROWSER/RUNTIME: browser name and version, operating environment, and relevant local runtime version]. Do not invent versions. The context matters because a local development server, test runner and browser can exercise different code paths, but the contract should contain only details relevant to reproducing this case.

Verification means repeating the refresh method without changing other variables and noting whether the observation is stable, intermittent or blocked. Do not manufacture a repeat count or success rate. An intermittent observation must remain labelled intermittent; it must not be rewritten as deterministic merely to make a cleaner prompt.

Startup command, revision and non-goals

Startup command is the repository-authorised command that starts the relevant development surface. Use [STARTUP COMMAND: supply the command from repository documentation or an authorised maintainer]. The fictional scenario does not imply a package manager, framework, container system or test runner. If dependencies are missing, report the block and ask before installing anything or using the network.

Revision pins the code under observation. Record [REVISION: branch or worktree identifier plus immutable revision, according to repository policy]. This prevents a later patch or unrelated update from being confused with the baseline. Do not create a branch or worktree unless the human has authorised that action.

Non-goals name behaviour that this investigation must not change. A suitable illustrative list is: no change to unrelated settings, public interfaces, authorisation behaviour, schema, dependencies, build scripts, translations or user-facing copy. Add repository-specific exclusions. Non-goals differ from unknowns: a non-goal is intentionally outside the task, whereas an unknown may need resolution to understand the defect.

Verify the startup command against repository-owned documentation and capture its status as run, failed, blocked or not run. Verify the revision before and after observation to detect unrelated changes. Stop if the documented command would mutate external data, require credentials not approved for this task or target an uncertain environment. A quick reproduction is not worth contaminating evidence or crossing an authorisation boundary.

Assemble a contract that Codex can restate without filling gaps

The completed packet should be machine-readable enough to follow and human-readable enough to review. Do not optimise it into a terse prompt that hides uncertainty. OpenAI’s prompting structure provides a useful distinction: desired behaviour says what should happen; reproduction says what was done; constraints limit the task; verification says how a change would later be checked. It does not guarantee diagnosis or correctness.

Illustrative contract template

The following is a recommended template, not a product feature or guarantee. Replace every bracketed field with authorised evidence or UNKNOWN:

Case status: fictional example adapted to an authorised local development copy. Route: [ROUTE]. Authorised test actor: [ACTOR OR FIXTURE]. Starting persisted value: [VALUE AND METHOD THAT ESTABLISHED IT]. UI action: [EXACT CONTROL CHANGE AND SAVE ACTION]. Success signal: [EXACT OBSERVED SIGNAL].

Refresh/reopen method: [EXACT METHOD]. Observed value: [VALUE AFTER THAT METHOD]. Expected value: [REQUIRED DURABLE VALUE AND BASIS]. Startup command: [REPOSITORY-AUTHORISED COMMAND]. Browser/runtime: [ACTUAL CONTEXT]. Revision: [PINNED REVISION]. Non-goals: [EXPLICIT EXCLUSIONS].

Safety boundaries: use test data only; do not expose secrets or untrusted private data; do not install packages, access the network, call external services or mutate a database without explicit approval. Do not edit files during observation. Do not take repository actions. List any requested capability that the current sandbox or approval policy does not allow.

Ask a human to read this contract against the original report or requirement. The reviewer should confirm the expected behaviour, environment, actor permission and non-goals. For consequential decisions—including changing access rules, schemas, external data or merge readiness—human review is required. Codex may organise evidence, but its restatement is not approval.

The acceptance rule for this contract is not “all fields contain text”. Each field must be evidence-backed, explicitly unknown or intentionally not applicable with a reason. Reject vague entries such as “normal user”, “latest code”, “refresh normally” or “run the app”; they cannot support repeatable observation.

Resolve contradictions before investigation

Contracts often expose contradictions. The starting value might be described as persisted even though it was observed only in the interface. The expected value might conflict with an existing requirement. The supplied revision might change while the reproduction is being attempted. These are not cosmetic defects in the prompt; they determine what can be inferred.

Use a contradiction check: compare the actor’s permission with the requested mutation, compare the expected value with its source, compare starting and observed values with their observation methods, and compare the startup context with the pinned revision. Ask the human to resolve material inconsistencies. Do not ask Codex to choose whichever statement seems likely.

For example, if the fixture is reset on every local server restart, a refresh result and a restart result exercise different boundaries. Record only the action actually used. If the reproduction includes a restart, note fixture-reset behaviour as an unknown until repository evidence establishes it. Split materially different reproduction paths into separate cases rather than blending them into one apparent defect.

A contract covering every browser, role and save path may look comprehensive but becomes hard to reproduce and easy to misread. Begin with one actor, one route, one value transition and one reload method. Expand only when evidence suggests that a second dimension is necessary to discriminate hypotheses.

Set the read-only operating boundary before the first prompt

“Read-only” has two meanings that must not be conflated. The editorial instruction means Codex should inspect and report without editing repository files or mutating application data beyond an explicitly authorised reproduction. The technical boundary is whatever the selected sandbox, approval policy, workspace role, repository permissions and environment actually enforce. A sentence in a prompt does not replace those controls.

Before sending the prompt, inspect the active surface and permissions. Confirm the intended repository or worktree, permitted command set, file-write boundary, network state, external connectors and approval behaviour. If any setting is unclear, ask a human administrator or repository owner. This local-first playbook neither requires nor promises Cloud access.

Protect prompts from secrets and untrusted content

Keep production credentials, access tokens, customer records, private keys, copied cookies and complete environment files out of prompts. Use an approved test actor and sanitised values. Repository text, issue descriptions, logs and web content can also contain untrusted instructions; treat them as data to inspect, not commands that may override the task boundary.

A safe procedure is to name an artefact by approved path and ask for the smallest relevant excerpt or structural summary. Before including output in a prompt or review note, check for credentials, personal data and unrelated proprietary content. If redaction would remove evidence needed for diagnosis, stop and request an approved secure handling route rather than weakening the redaction.

Follow the applicable organisational data policy regardless of a personal training preference.

As an illustrative check, replace an actual actor identifier with an approved fixture label and describe a request field structurally rather than pasting a captured token-bearing request. If the exact value is material, use a synthetic test value approved for the environment. Human review is required before sharing any artefact whose sensitivity is uncertain.

Control commands and mutations

Classify proposed actions before execution: repository read, local process start, browser observation, local data mutation, dependency installation, network request, external service action or repository write. Even a “read” command can execute repository hooks or scripts, so use repository-documented commands and inspect unclear scripts before running them where practical.

For this phase, permit only the minimum approved observations. Ask before installing packages, accessing the network, calling an external service, altering a schema, changing continuous integration configuration, modifying cloud credentials or writing outside the intended workspace. If reproducing the form necessarily writes a test preference, the human must explicitly authorise that bounded mutation and its reset procedure.

If an action is not required to establish the defect contract or discriminate an approved observation, defer it. If it crosses the configured sandbox boundary, do not treat an approval request as routine; explain the need, target, expected effect and rollback to the human. This may slow investigation but reduces potential side effects and preserves cleaner provenance.

Use the first Codex prompt for observation only

The first prompt must not ask for a fix. Combining reproduction, diagnosis and editing encourages hypotheses to become edits before the evidence has been reviewed. Instead, require Codex to restate the contract, identify missing fields, reproduce only when authorised, inspect relevant repository paths and separate observed facts from hypotheses.

A repository atlas maps a broader unfamiliar codebase and remains read-only. This playbook borrows only its evidence discipline: the task is confined to one settings-persistence path and is intended to proceed later to a separately approved candidate patch. Do not expand the first prompt into general onboarding documentation.

Copy-ready read-only prompt

The following is a recommended example. Replace all placeholders, remove unsupported permissions and have a human review the packet before consequential execution:

Role and task: “Investigate one authorised, fictional-example settings-persistence reproduction in read-only mode. Do not edit, create, delete, rename or format files. Do not install dependencies, use the network, call external services or change application data except for the explicitly authorised test-fixture action below. Do not commit, push, open a pull request, or merge.”

Defect contract: “Route: [ROUTE]. Authorised test actor: [ACTOR]. Starting persisted value and establishment method: [STARTING VALUE AND METHOD]. UI action: [ACTION]. Success signal: [SIGNAL]. Refresh/reopen method: [METHOD]. Observed value: [OBSERVED]. Expected value and requirement source: [EXPECTED]. Startup command: [COMMAND]. Browser/runtime: [CONTEXT]. Revision: [REVISION]. Non-goals: [NON-GOALS].”

Required procedure: “First restate the contract without resolving omissions yourself. List contradictions, missing evidence and actions requiring approval. If the authorised environment and command are available, attempt only the stated reproduction. Otherwise report it as blocked or not run; do not substitute another environment or command.”

Repository observation: “Inspect only files relevant to the control’s rendered value, validation, submit path, success signal and reload path. Cite repository-relative paths and line ranges where available. Identify existing relevant tests and repository-documented verification commands, but do not run a command unless it falls within the approved read-only observation scope.”

Evidence discipline: “Report observed facts separately from hypotheses. For every observed fact, state its source: direct UI observation, approved command, source reference or permitted runtime observation. Do not infer a backend, endpoint, payload, storage system, framework or root cause. List all unknowns explicitly.”

Output: “Return: contract restatement; omissions and contradictions; actions performed; observations with evidence references; relevant files and their apparent roles; existing tests or commands found; at least two competing hypotheses only where evidence supports them; the smallest non-destructive observation that could distinguish each; blocked items; and questions for the human owner. Propose no patch and make no edits.”

Verify the response against the working tree and execution record. Check that no file changed, no unapproved process or network action occurred, every purported fact has a source and every uncertainty remains labelled. Do not accept a confident narrative as a substitute for evidence. If the tool reports an action that cannot be corroborated, classify it as unverified and ask for clarification.

The pass rule for this first prompt is procedural, not diagnostic: the contract is accurately restated, observations and hypotheses are separated, unknowns are listed, relevant paths are cited where available and no prohibited action occurred. Finding a root cause is neither required nor expected. If the response proposes edits anyway, discard the proposal from the evidence record and reissue the boundary more narrowly.

Build an observation ledger without claiming a root cause

An observation ledger turns a narrative response into reviewable units. Each entry should include an identifier, proposition, evidence type, source location, observation conditions, confidence limitation and status. This differs from a hypothesis list: ledger entries say what was seen or read, while hypotheses explain how several observations might fit together.

For example, a permissible illustrative entry is: O-01 — After the authorised Save action, the configured success signal appeared. Evidence: direct UI observation at the pinned revision. Limitation: no durable-state conclusion. Another could be: O-02 — After the specified refresh, the control rendered the starting value. Evidence: direct UI observation. Limitation: the source of the reloaded value remains unknown. These examples are formats, not claimed results.

A source-inspection entry might state: O-03 — A named repository file appears to define the control’s initial rendered value; cite path and lines. Limitation: static source inspection does not confirm that this branch executed. Do not include fictional paths or line numbers. Codex must either cite actual authorised repository locations or write not located.

Use status labels that preserve failure and uncertainty

Give every planned observation one of four statuses: observed, failed, blocked or not run. “Failed” means the attempted procedure produced an actual failure; “blocked” means a prerequisite or permission prevented it; “not run” means it was deliberately deferred. Do not collapse the latter three into “unable to verify”, because they imply different next decisions.

For example, a startup command that exits unsuccessfully is failed, while a command withheld because dependencies would require network installation is blocked pending approval. A browser repeat omitted because the authorised actor is unavailable is not run. None of those statuses permits Codex to claim that the defect reproduced.

A human must compare status labels with the command or observation record, using only safe excerpts. Do not paste entire logs when a minimal redacted excerpt and command status suffice. Absent evidence remains absent: never convert “not run” into an assumed pass, and never interpret “blocked” as evidence for or against a hypothesis.

Keep hypotheses competitive

Once observations exist, Codex may list plausible explanations, but it should not nominate a winner without a discriminator. Relevant categories could include a value held only in client state, a field absent from submission, rejected validation, a write that targets an unintended record, a stale reload or cache, an incorrect read mapping, or a fixture/environment reset. These categories are examples, not assertions about the fictional application.

For each hypothesis, ask for three fields: evidence consistent with it, evidence that would count against it, and the smallest non-destructive observation capable of distinguishing it. For an omitted-field hypothesis, the discriminator may involve a permitted inspection of how the submitted value is constructed. For a stale-read hypothesis, it may involve comparing the approved durable read with the rendered reload value. The actual method depends on the repository and authorisation boundary.

Reject hypotheses that merely restate the symptom, such as “persistence is broken”. Also reject explanations that introduce an unseen endpoint, database or cache. If only one hypothesis is listed despite several remaining possibilities, ask Codex to identify alternatives compatible with the evidence. Conversely, do not demand an arbitrary number when the repository evidence genuinely narrows the field; diversity is useful only when each candidate is plausible.

Proceed from observation to patch planning only when one explanation has materially stronger evidence or when a bounded change can address the demonstrated fault without relying on an unverified implementation story. If discriminating observations require production access, broad network permission or sensitive data, stop and escalate to a human owner rather than broadening Codex’s access.

Systematic isolation is useful when competing explanations remain live, but this case retains a narrower contract: one UI action, one persistence boundary and one later candidate patch. Use broader debugging techniques only to resolve a stated unknown, not to transform this section into an open-ended hunt across the codebase.

Review the read-only packet before allowing any write phase

The observation phase ends with a human checkpoint, not with Codex declaring that it understands the bug. The reviewer should receive the frozen contract, action list, observation ledger, repository references, competing hypotheses, blocked items and proposed discriminators. This is an investigation packet, not a merge packet and not evidence that a patch exists.

Apply a five-part readiness decision

First, check identity: the route, actor, environment and revision in the observations must match the contract. Second, check reproducibility: the exact action and reload boundary must be documented, with status rather than an invented outcome. Third, check provenance: each material claim must point to an authorised observation or repository reference. Fourth, check uncertainty: hypotheses and unknowns must remain visibly distinct from facts. Fifth, check scope: no proposed next step should alter non-goals or require undeclared access.

An illustrative review decision could be: “Ready for one additional read-only discriminator because the submit construction is identified but the reload source is unknown.” Another could be: “Not ready for planning because the starting value was never established as persisted.” These are sample decision formats, not outcomes. The reviewer should name the missing evidence and the least invasive way to obtain it.

The go/no-go rule is conservative. Move towards a plan-before-write phase only if the defect contract is complete enough to judge a proposed change, the observation record supports a bounded fault area, and unresolved risks are explicit. Stay read-only if material implementation paths remain unknown. Stop entirely if authorisation, data handling or environment ownership is uncertain.

Editing sooner may produce a plausible diff quickly, but it makes it harder to distinguish a genuine repair from a change that merely alters the observed symptom. A reviewed observation packet costs time up front and gives the later regression design a defined persistence boundary.

Record limitations that must survive into later packets

Carry forward every unresolved limitation. If no permitted runtime trace was available, say so. If only static source inspection was performed, say so. If the browser reproduction was blocked, do not let a later test substitute silently for it. If the starting persisted value was inferred rather than read, resolve that weakness before claiming durable-state verification.

Also preserve the non-goals verbatim unless a human explicitly changes scope. A later proposal must not quietly alter interfaces, schema, dependencies, permissions, unrelated settings, build scripts, translations or copy merely because those changes appear convenient. Scope changes require a reason, impact analysis and renewed human approval.

OpenAI’s prompting guidance says that, after a fix, Codex should repeat the reproduction and, where relevant, report commands and results for lint and the smallest relevant test suite. That later procedure must report actual outcomes as pass, fail, blocked or skipped; this read-only phase claims none. A successful command will remain distinct from durable-state evidence and from the repeated browser path.

Model or product-surface availability should not become part of the defect theory. Use only the authorised surface available to the account and workspace.

Conclude the read-only phase with a named human decision: continue observation, approve preparation of a scoped plan, revise the contract or stop. Do not ask Codex to approve its own evidence. No consequential repository, permission, data or merge decision should proceed without human review, and no instruction in this playbook authorises a commit, push, pull request or merge.

Map the full persistence path before choosing a patch

The fictional defect is narrow: in an authorised development copy, a reader visits /settings/notifications, changes Enable alerts, selects Save, sees a success signal, refreshes or reopens the view, and finds the old value rendered again. That sequence establishes a user-visible contradiction, not its cause. The investigation must therefore trace two related but distinct paths: the write path that should carry the changed value towards durable state, and the read path that should retrieve that state after reload.

Ask for a persistence map that names observed behaviour, relevant code paths, constraints and a verification method, rather than asking Codex to “find and fix the bug”. The former request creates reviewable intermediate evidence; the latter encourages a premature conclusion. This method is an editorial recommendation, not a guarantee that Codex will identify the root cause or propose a correct patch.

Do not assume the application uses a representational state transfer (REST)An architectural style for web application interfaces; REST-based interfaces use resource representations and standard web request methods. Open glossary entry endpoint, a database, a particular front-end framework, a server at all, or even a conventional form submission. The actual mechanism could be a command, remote procedure call, local process, event, queued write, generated client, browser-backed store or another repository-specific boundary. The map should use neutral functional stages until repository evidence supplies concrete names.

Define the map as observable stages, not architectural guesses

Begin with the control’s displayed state and follow the value through every stage that can transform, omit, reject, cache, write or read it. A useful neutral sequence is: rendered control value; control-state update; validation or normalisation; submit initiation; submitted representation; client cache or store update; permitted request observation; receiving handler; server-side validation or authorisation; write operation; write acknowledgement; reload query; response mapping; cache reconciliation; and rendered value after reload.

That sequence is a checklist, not a claim that all stages exist separately. For example, validation may occur inside the submit handler, and there may be no client store. Conversely, one visible control may map to several fields or commands. Mark each stage as observed, inferred, absent by evidence or unknown, then attach the repository path, symbol, permitted runtime observation or test that supports the classification.

A stage may receive a concrete architectural label only when authorised evidence identifies it. If a file appears to declare a route but runtime evidence is unavailable, record “route declaration observed; invocation unverified”. If a browser observation shows an outbound operation but the receiving implementation is outside the workspace, record that boundary as unknown. Do not fill the gap with a likely framework convention.

Use the following illustrative prompt packet only after replacing bracketed fields with authorised evidence. It requests analysis without edits and deliberately separates file inspection from runtime observation:

Persistence-path mapping packet — read only

Case:
Route or entry point: [authorised route]
Control: [control label and stable selector if known]
Starting durable value: [verified test value or UNKNOWN]
Attempted new value: [test value]
Save signal: [observed signal]
Reload method: [exact refresh, remount or reopen sequence]
Observed post-reload value: [observed value]
Expected post-reload value: [expected value]
Revision and environment: [authorised details]

Do not edit files, install packages, change configuration, create commits,
push, open a pull request, alter data outside the authorised test fixture,
or make external calls.

Trace this value through these functional stages where they exist:
1. rendered control state;
2. state update caused by the interaction;
3. validation, coercion or normalisation;
4. submit handler;
5. submitted representation;
6. client cache or store;
7. permitted request or command observation;
8. receiving handler;
9. server-side validation, authorisation and write;
10. acknowledgement or success-state decision;
11. reload query or durable-state retrieval;
12. response-to-view mapping;
13. cache reconciliation;
14. rendered value after reload.

For every stage, provide:
- status: OBSERVED, INFERRED, ABSENT-BY-EVIDENCE or UNKNOWN;
- repository path and symbol, with line range where stable;
- value name and any transformation;
- evidence source;
- uncertainty and the next non-destructive observation.

Do not assume REST, a database, a framework, an endpoint, a payload shape
or a persistence technology. Keep facts separate from hypotheses. Stop
without proposing or applying a patch.

Review the response line by line. A path citation supports what the code says, but does not prove that the same branch ran during reproduction. A runtime observation can show that an operation occurred, but does not necessarily reveal the implementation that handled it. A test can document intended behaviour, but may use a seam that bypasses the defect. A compact map with explicit unknowns is more useful than a comprehensive-looking map built from assumptions.

Conceptual illustration of a value travelling through abstract client state, request, durable store and reload nodes, with one ambiguous break point and no words or real architecture.
The investigation follows the value across state, write and read boundaries.

Trace the control state without mistaking it for persisted state

At the first stage, ask where the displayed value originates on initial render and what changes when the user toggles it. The answer should distinguish a component-local value, a form abstraction, a shared client store, a loader result and a server-derived value only where those categories are demonstrated by code. A control changing visually proves that an interaction updated some rendered state; it does not prove that the submitted representation includes the change.

For the fictional example, a valid sample entry might read: “The control is initialised from symbol [preference source] in [path]; the interaction calls [handler]; whether that handler supplies the save operation is unknown pending inspection.” This is a template, not a finding about a real repository. An invalid entry would say “the React state updates correctly” when neither React nor the relevant execution has been established.

Identify the initial-value expression, the interaction handler and every transformation between them. Check for inversion, defaulting, string-to-Boolean conversion, disabled controls, conditional registration, stale closure capture and name mismatches, but report these only as inspected possibilities. Compare the value immediately before interaction, immediately after interaction and at the submit boundary using existing safe tests or permitted local observation; do not introduce logging until a human approves an edit.

If the control-state transition itself is unverified, stop the downstream diagnosis from being described as complete. You may continue reading code to understand candidates, but the map must retain the gap. Instrumenting the control could provide stronger evidence, yet instrumentation is a write and may change timing or behaviour. Prefer existing debugger facilities, tests or development tooling when authorised; otherwise request approval for the smallest temporary instrument and a removal plan.

Inspect validation and submission as separate gates

A changed control can be lost before submission. Validation may reject it, normalisation may restore a default, a dirty-field filter may omit it, or a submit handler may read a different source from the one rendered. These are competing mechanisms, not a proposed diagnosis. The map should show the value entering and leaving each demonstrated gate, including the branch that decides whether the success signal appears.

Ask Codex to answer concrete questions: Which event invokes submission? Is submission prevented or redirected? Which symbols build the submitted representation? Does validation return an error, a transformed value or no value? Does the success indicator depend on confirmed durable acknowledgement, mere handler completion, an optimistic update or some other evidenced condition? If this cannot be determined from authorised evidence, require “unknown” rather than an inferred answer.

An illustrative mapping row could say: “The save action calls [submit symbol]. The representation builder reads [source A], while the visible control reads [source B]. Their equivalence has not been demonstrated.” That row supports a hypothesis about divergence without declaring a defect. The discriminator would be a non-destructive observation of both values during one authorised reproduction, not an immediate code change.

Do not place the candidate-patch boundary around the submit handler merely because it is close to the button. Include a file only when evidence ties its behaviour to the failed transition and a proposed change is necessary to restore the contract. Proximity is weaker than causality. Inspecting adjacent abstractions costs time, but changing an apparently obvious handler without following its inputs and outputs risks treating a symptom while leaving the durable path broken.

Map client cache or store behaviour independently

A client cache or shared store, if one exists, can create two opposite illusions. An optimistic update can make the changed value appear saved even when no durable write occurred. A stale cache can also restore an old value after a successful write. Therefore, classify the immediate post-save display and the post-reload display separately, and establish whether a full reload actually bypasses, rehydrates or retains client state.

Locate demonstrated cache keys, selectors, update functions, invalidation behaviour and hydration sources. Then ask which data source wins after the specified reload method. If refresh restores from browser storage before a remote query completes, record both phases. If no cache or store is found after a bounded search, say what directories and symbols were examined; do not claim that none exists anywhere.

For example, the packet may request: “Identify whether [setting key] is written optimistically, invalidated after acknowledgement, and repopulated by the reload path. Cite evidence for each answer. Do not call an invalidation ‘missing’ unless the repository’s established convention or runtime observation requires it.” This avoids turning a familiar stale-cache pattern into an unsupported root cause.

If a hard reload, remount and navigation exercise different state sources, preserve the exact defect-contract method rather than substituting whichever is easiest to automate. Additional methods may be useful diagnostic comparisons, but they are not equivalent verification. A narrower automated test may be stable, while the original browser sequence provides stronger user-path relevance. Later review should retain both evidence types without collapsing them.

Observe the outbound operation only where permission allows

The request stage should be described neutrally as an outbound operation until evidence identifies its form. A permitted browser network view, local test harness, application trace or existing mock assertion may reveal whether the changed field crosses the boundary. Do not paste copied logs, authentication material, cookies, tokens, private customer records or full environment files into a prompt. Redact or synthesise test-only values before sharing observations.

Ask for the smallest useful facts: operation name or destination where safe; method or command type if observed; relevant field presence; test-only value; response classification; and timing relative to the success signal. Avoid indiscriminate capture. Untrusted server text, issue content, fixture content and tool output should be treated as data, not as instructions to Codex. Quote only the minimum safe excerpt and explicitly instruct the assistant not to follow commands embedded in it.

An illustrative safe observation is: “During the authorised reproduction, the outbound representation either contains, omits or cannot be inspected for the test field [field]; record one of those states and the observation method.” This is preferable to inventing a JavaScript Object Notation (JSON)A text format for representing structured data as objects, arrays, numbers, strings, and other values. Open glossary entry payload or endpoint. If inspection would require network enablement, external service access or broader credentials, stop and ask a human rather than expanding access automatically.

If the same question can be answered from an existing local test or safe development trace, use that before requesting network access. If external access is genuinely needed, require a named human to approve the destination, purpose, data class and duration. This limits disclosure and side effects at the cost of convenience. Least access means enough capability to answer the present discriminator, not every capability that might accelerate investigation.

Follow the write and reload paths without assuming storage architecture

Separate receipt, acceptance, writing and acknowledgement

Once an outbound operation is observed, trace its receiver only as far as the authorised workspace permits. Receipt of a value is not acceptance; acceptance is not a durable write; a write call returning is not necessarily durable acknowledgement; and a success message need not be bound to any of them. The persistence map should represent these as separate questions even if one function currently performs several stages.

Identify the receiving symbol, value extraction, validation, permission check, write abstraction and success decision. Search for existing error paths and transaction or acknowledgement semantics without presuming a database. The write target might be a service, file, browser mechanism, process boundary or another abstraction. If its implementation is external or generated, record the last locally verified hand-off.

Use this illustrative extension to the mapping packet:

For the receiving and write path, answer without edits:
- What evidence shows the operation reaches the receiver?
- Where is the relevant value extracted and under what name?
- What validation, normalisation or permission decision can alter it?
- What abstraction receives the write request?
- What event causes the caller to treat the save as successful?
- Does available evidence establish durable acknowledgement, or only
  acceptance/dispatch/completion at an earlier boundary?
- What remains outside this workspace or permission scope?

Do not name a database, transaction, public request schema or authorisation defect
unless repository or authorised runtime evidence establishes it.

If the evidence stops at dispatch, phrase the hypothesis as “the value may be lost at or beyond the write boundary”, not “the database failed to save”. That language preserves the scope of knowledge. It prevents a candidate patch from reaching into schema, permissions or infrastructure without evidence and approval.

Trace the reload query as an independent source of failure

The write path may be correct while the reload path reads the wrong key, stale projection, different scope or mismapped representation. Consequently, begin the read path at the exact refresh or reopen action from the defect contract. Identify what initiates retrieval, which actor or test fixture scopes it, what value returns at the nearest observable boundary, how that value is mapped, and which source ultimately renders the control.

A useful procedure compares three values without claiming any result in advance: the attempted value at submission; the value returned by the write acknowledgement, if one exists; and the value supplied to the control after reload. If these differ, mark the first demonstrated divergence. If they agree but the display differs, inspect mapping and rendering. If no durable-state retrieval can be observed, state that the regression oracle remains incomplete.

For the fictional case, a sample request is: “Trace the post-refresh source for Enable alerts. Determine whether the view reads the same setting identity and actor scope that the save path writes. Cite the symbols that establish sameness; otherwise record the relationship as unresolved.” This wording avoids presuming that the setting is a column, preference record or endpoint field.

A patch aimed at writing is not justified when the earliest evidenced divergence is on reading, and a cache invalidation patch is not justified when the durable retrieval itself returns the old value. Where observation cannot locate the divergence, do not edit either side. Obtaining read-boundary evidence may require a more involved local fixture, but it materially reduces the risk of patching the wrong layer.

Require a stage-by-stage evidence matrix

Ask Codex to turn the narrative trace into a compact matrix with columns for stage, input value identity, output value identity, evidence, certainty, and next discriminator. Value identity matters because names can change legitimately across boundaries. The matrix should state how equivalence is established—for example, through an explicit mapping function—rather than assuming similarly named fields are the same.

A clearly labelled illustrative row set might contain: “control state: observed in [path/symbol]”; “submitted representation: unknown”; “receiver extraction: inferred from declaration, execution unverified”; “write abstraction: observed statically”; “reload retrieval: unknown”; and “rendered post-reload value: observed in authorised reproduction”. These placeholders demonstrate reporting form only; they do not describe actual files or outcomes.

Verification consists of challenging every causal arrow. Ask, “What evidence shows this output becomes that input?” A call graph supports reachability; one reproduction supports execution; a fixture supports controlled state; and a returned value supports a boundary observation. None alone proves the entire chain. Require Codex to mark conflicting evidence rather than resolving it by preference.

The map is ready for hypothesis testing when it identifies at least one bounded divergence or two plausible divergence points with non-destructive discriminators. It is not ready merely because every box has text. Stop exploring unrelated subsystems once the competing explanations can be distinguished within the approved scope.

Keep at least two hypotheses alive until evidence discriminates

Construct hypotheses from observed divergence, not pattern matching

A useful hypothesis states a mechanism, the evidence that supports it, evidence that would weaken it, and the smallest non-destructive discriminator. Require at least two evidence-supported hypotheses whenever the observations permit more than one explanation. “The save is broken” is not a hypothesis because it neither locates a mechanism nor predicts a distinguishing observation.

Potential categories include client-only success feedback, omission from the submitted representation, validation rejection or coercion, stale cache reconciliation, write-path mismatch, read-path mismatch, and fixture or environment inconsistency. These categories are prompts for inspection, not findings. Codex should include one only when a cited observation makes it plausible, and it should not force two artificial hypotheses when evidence has already disproved alternatives. In that exceptional case, require the response to show the disproof.

Use the following packet after the persistence matrix has been human-reviewed:

Competing-hypotheses packet — no edits

Using only the approved defect contract and cited persistence-path evidence,
provide at least two plausible hypotheses where the evidence supports them.

For each hypothesis list:
1. precise mechanism;
2. supporting observations, with paths/symbols or permitted runtime evidence;
3. observations that conflict with it;
4. predicted observation if it is true;
5. predicted observation if it is false;
6. the smallest non-destructive discriminator;
7. access or approval needed;
8. residual uncertainty after that discriminator.

Rank hypotheses by evidence strength, not familiarity. Do not infer a
framework, endpoint, database, schema or root cause. Do not modify files,
fixtures, settings, dependencies or external systems. If fewer than two
hypotheses remain credible, show the evidence that eliminated the others.

Human review should reject a hypothesis that merely restates a code smell. For instance, “cache invalidation is missing” is unsupported unless the application has an evidenced cache, the relevant convention requires invalidation, and the reload sequence can consume stale cached state. Retain a hypothesis only if it predicts an observation that differs from at least one competitor.

Choose minimal, non-destructive discriminators

A discriminator answers one contested question while changing as little as possible. Examples include inspecting an existing test assertion for field inclusion, pausing at an already available debugger breakpoint, comparing a test-only value at two local boundaries, or repeating the reproduction with the same authorised fixture after resetting it through an established safe procedure. These are suggested methods, not guarantees that the repository supports them.

Consider two illustrative hypotheses. Hypothesis A says the visible control updates one state source while submission reads another, so the outbound representation lacks the changed value. Hypothesis B says the submitted value crosses the boundary, but the reload mapping reads a different setting identity. The smallest discriminator for A is an authorised observation at the submit boundary. If the changed value is present, A weakens and the investigation moves downstream. If absent, B has not been disproved, but A gains support.

A second observation may then discriminate write from read behaviour: inspect the value returned through an existing local durable-state seam after save and before rendering. Do not create or mutate a production-like record merely to answer this question. Use a dedicated test fixture or authorised non-production account, and reset it according to repository practice. If no safe seam exists, record the block and request human guidance.

Prefer the discriminator with the least mutation, narrowest data exposure and clearest ability to separate hypotheses. Do not run a broad test suite, enable unrestricted network access or alter storage merely because those actions might reveal more. Several small observations may take longer than one broad experiment, but they preserve attribution and reduce collateral effects.

Record outcomes without manufacturing certainty

Each discriminator should produce one of four states: supports, weakens, inconclusive or blocked. “Supports” is not “proves”. An observation may be consistent with several mechanisms, and static code may differ from the executed branch. Record the exact environment, revision, fixture identity in non-sensitive form, method and limitation so another authorised reviewer can judge whether the inference is warranted.

An illustrative outcome format is: “Question: does the changed value reach [boundary]? Method: [authorised observation]. Status: not yet run. Possible interpretations: present narrows attention downstream; absent narrows attention upstream; unavailable leaves both hypotheses open.” Keeping the sample unexecuted avoids inventing results while showing the required reasoning structure.

A human must compare the recorded observation with the predicted true and false states. If the observation was not one of the predictions, revise the hypotheses rather than forcing it into either. If environment or fixture differences could explain the result, add that as a competing hypothesis and seek a minimal consistency check.

No candidate patch should be proposed until at least one hypothesis is materially stronger than its competitors and the earliest evidenced divergence lies inside the authorised repository scope. If uncertainty remains evenly split across layers, continue read-only investigation or escalate. Delayed editing avoids a speculative multi-layer patch that obscures which change matters.

Draw the candidate-patch boundary before any edit

State what cannot change by default

The candidate patch is not “everything near settings”. Its default boundary excludes any API or message-contract shape change, schema or migration change, permission or authorisation change, unrelated settings behaviour, dependency addition or upgrade, build-script modification, translation or copy change, and public interface change. It also excludes production data, cloud configuration, continuous integration configuration, repository policy, generated artefacts not governed by existing practice, and external services.

These exclusions are not claims that such layers can never contain the defect. They are decision gates. If evidence points beyond the boundary, Codex should stop and prepare a scope-change request rather than silently widening the patch. The request should identify the newly implicated layer, evidence, proposed effect, alternatives considered, additional reviewers and rollback implications. A human then decides whether to authorise a new investigation boundary.

For example, if the strongest hypothesis requires changing a shared public setting type, the response should not disguise that as a local form fix. It should say that the proposed correction affects a public interface and therefore violates the current boundary. The practical options are to find a local adaptation supported by conventions, obtain explicit scope expansion, or leave the issue unresolved. Convenience never overrides an explicit non-goal.

Apply the smallest-file and smallest-behaviour tests

For each proposed path, ask two questions: what evidenced defect mechanism requires this file to change, and what would remain broken if it did not change? A file fails the boundary test if the rationale is merely “related”, “cleanup”, “consistency” or “adjacent”. Tests may require a separate path, but their purpose must be to exercise the durable-state boundary or an established supporting seam, not to refactor unrelated coverage.

The smallest-behaviour test asks whether the proposal restores only the contract: changing Enable alerts through the authorised route should survive the specified reload for the same authorised actor. It must preserve other setting semantics, validation, permissions, request contracts and public behaviour unless a human deliberately broadens scope. A patch that rewrites the settings subsystem may use fewer conceptual abstractions but still have a much larger behavioural surface.

An illustrative boundary table should list “candidate path”, “necessary mechanism”, “expected behaviour change”, “behaviour preserved”, “risk”, and “why no adjacent path changes”. Empty rationales are blockers. Verification is a human comparison against repository ownership and policy, because Codex cannot supply organisational approval merely by presenting a coherent plan.

Choose the proposal with the narrowest evidenced behavioural effect, not automatically the fewest changed lines. A slightly larger explicit mapping correction may be safer than a one-line global default change that affects every setting. State the consequence: local duplication or adaptation can constrain impact, whereas changing a shared abstraction may improve consistency but demand broader tests and reviewers.

Demand a plan-before-edit response

The final output of this section is a proposed plan, not a patch. It must list every path expected to change, the precise rationale, the hypothesis addressed, risks, rollback method and unresolved uncertainty. It should also list inspected paths that will not change and explain why. No file modification should occur until an authorised human accepts the boundary under the repository’s own process.

Scoped candidate-patch plan — stop before editing

Using the reviewed evidence matrix and discriminator record, propose the
smallest candidate patch. Do not edit any file.

Current leading hypothesis:
[evidence-supported statement]

Contract to restore:
[exact authorised save-and-reload behaviour]

Hard boundaries unless a human changes scope:
- no API or message-contract shape change;
- no schema, migration or persistence-structure change;
- no permission or authorisation change;
- no unrelated settings behaviour change;
- no dependency addition, removal or upgrade;
- no build-script or CI change;
- no translation or copy change;
- no public-interface change;
- no production data, external action, commit, push, PR or merge.

Return:
1. each path proposed for change;
2. symbol or region involved;
3. evidence-based rationale;
4. why adjacent paths do not need changes;
5. expected behaviour after the candidate patch;
6. regression evidence to add or update, without claiming a result;
7. risks and possible collateral effects;
8. rollback procedure;
9. unresolved uncertainties;
10. permissions or approvals needed;
11. explicit confirmation that no edits were made.

If evidence requires crossing a hard boundary, provide a scope-change request
instead of a patch plan. Stop for named human approval.

Review the plan for hidden widening. “Update types” may imply a public contract; “refresh dependencies” violates the dependency boundary; “fix translations” introduces unrelated copy work; “adjust test setup” may alter shared fixtures or build scripts. Require exact paths and effects before approval. If the repository path is not yet known, the item remains unresolved and cannot be authorised as an unspecified wildcard.

Rollback must match the proposed change. Reverting a candidate source edit may be sufficient for a local mapping correction; it is not necessarily a credible rollback for a schema, permission or external-state mutation. For that reason, those changes are outside this boundary. The plan should also preserve the pre-edit reproduction record so reviewers can compare later evidence against the same contract.

Approve editing only if the leading hypothesis is evidence-supported, all changed paths have necessary rationales, the regression approach crosses the relevant persistence or reload boundary, rollback is feasible, and unresolved uncertainty is proportionate. Otherwise return to mapping or discriminator work. Human approval is consequential here because the next phase changes repository state and may affect shared behaviour.

Make least access and approval operational

For this phase, keep work read-only until the plan is approved. When edits are authorised later, limit writes to the intended repository or worktree and approved paths. Ask separately before package installation, network access, external calls, data mutation outside the dedicated fixture, writing outside the task branch, pushing, opening a pull request, changing continuous integration, or altering credentials. A prompt saying “do not push” is an instruction; enforceable limits depend on actual sandbox, approval, role, network and repository controls.

Use test-only data and keep secrets and untrusted content out of prompts. This local playbook neither requires nor promises Cloud, network tools, a particular model or identical limits.

The practical approval procedure is to name the reviewer, present the evidence matrix, hypothesis ranking, proposed paths, tests to be attempted, requested capabilities and rollback, then record approve, revise or reject. Approval should be limited to the stated revision and scope. Any new file, external action, schema implication, permission implication or unexplained test requirement returns the task to review.

Use the least capability that can perform the approved next step, and require human review before repository edits, scope expansion and every consequential external action. Even after approval, Codex must not merge, push or open a pull request under this playbook. The output of this section is a reviewable candidate-patch plan; whether work proceeds is a human governance decision, not a product guarantee.

Build a regression oracle around durable state, not the success signal

The candidate patch is not ready for review merely because the fictional Enable alerts control stays selected immediately after Save. The regression oracle must cross the same durable-state boundary implicated by the defect contract: set a different authorised test value, invoke the real save path, discard or bypass the current view’s transient state, read the value again through the permitted persisted source, and compare that returned value with the intended value. This distinguishes a persistence regression from a component-state check. Identify the write action and the independent read action before choosing an assertion. For example, in the authorised development copy of /settings/notifications, the proposed test could start with alerts disabled, save alerts as enabled, remount or reload the settings view, and then assert that the value supplied by the post-reload read path is enabled. This sequence is illustrative; the repository’s actual route, test fixture, state source and conventions must replace every placeholder.

For a durable-state regression, make the verification method explicit: “prove the changed preference survives the relevant write/read boundary”, rather than “add a test for Save”. If a proposed test can pass when no durable write occurs, it is not the primary regression for this defect. It may still be useful as a narrower unit check, but it cannot carry the persistence claim.

Define the oracle before discussing test syntax

An oracle states what observation decides whether the repaired behaviour is present. It is separate from the test runner, framework or assertion library. Because those repository details are unknown, do not ask Codex to invent a command, endpoint, database or helper. Instead, supply an evidence-backed behavioural sequence and ask it to map that sequence onto existing test conventions. A recommended procedure is: record the initial durable value; choose a different test value; perform the supported save operation; confirm only that the operation completed far enough to permit a read; invalidate the current presentation state by reload, remount or a permitted direct persisted-source query; retrieve the value again; and compare it with the chosen value.

For example, an oracle description might say: “Given the authorised fixture’s durable notification preference is false, set it to true through the existing settings submission path. Then create a fresh view instance or use the existing repository helper that reads the persisted preference. Assert that the fresh read yields true. Restore or dispose of the fixture according to existing test conventions.” This is a proposed method, not a claim about available helpers or observed results. The meaningful trade-off is fidelity against cost and isolation. A fresh browser reload may resemble the user’s path most closely but can be slower or less diagnostic; a permitted persisted-source query may be more precise but can bypass a faulty reload mapping. Choose the highest-fidelity boundary that is stable and authorised, then retain a smaller diagnostic check only if it explains failures without replacing the durable oracle.

Write this requirement into the prompt: accept an oracle only if the test’s post-save observation cannot be satisfied solely by the same in-memory object, optimistic cache entry or component instance that initiated the save. If the repository cannot provide such a seam without broad infrastructure work, report the regression as blocked or propose a proportionate integration test; do not weaken the assertion silently. A human reviewer must decide whether the absence of an automated durable-state test is acceptable for this consequential merge decision.

Reject the optimistic-state false positive

An optimistic user-interface update changes local presentation before, or independently of, confirmed durable storage. It can make an interaction feel immediate, but it cannot prove that a later session will read the new value. A test that clicks Save and immediately inspects the still-mounted toggle may therefore pass even if the write is omitted, rejected, directed at the wrong field or followed by an old-value read. The practical review procedure is to trace the value used by the final assertion. Ask: did it originate from the post-save component state, a client cache, a mocked response, an acknowledged write, or a fresh persisted read? Record the answer rather than treating all sources as equivalent.

Consider two clearly labelled illustrative checks. Check A mounts the settings form, selects Enable alerts, invokes the submit handler and expects the control to remain selected. Check B establishes a known starting preference, selects Enable alerts, performs the supported save, destroys the first view, creates a fresh read context and expects the reloaded value to be enabled. Check A may validate event wiring or optimistic rendering. Check B is the persistence candidate because it forces a separate read after the save. Neither check has actually run in this fictional case. Check A may accompany Check B but must not be reported as “persistence verified”. If only Check A is feasible, label the durable regression blocked and explain which missing fixture, permission or test seam prevents Check B.

A Saved message is similarly weak evidence. It can demonstrate that a success branch rendered, but it does not by itself establish what was written or what a future read returns. The recommended procedure is to treat the message as one intermediate observation in the repeated reproduction record, then continue through refresh or reopen. For example, record “success signal observed” in one field and “freshly read preference” in another. Asserting every implementation detail can make a test brittle, while asserting only the toast misses the defect. Prefer the smallest sequence that includes the durable write/read transition and the final user-relevant value.

Conceptual illustration of a small candidate patch passing through a regression loop and a human merge-review gate, no code, screenshot or readable letters.
A candidate patch needs durable-state evidence and a separate human merge decision.

Select the permitted way to cross the boundary

There are three broad, editorially recommended boundary-crossing patterns, but their availability must be established from the repository. First, reload or recreate the user interface so that it obtains settings through its ordinary read path. Second, remount the relevant view after clearing only the transient state that the application itself would discard. Third, query an authorised persisted source through an existing test helper, then separately test the read mapping where necessary. Do not invent a database table, network endpoint or browser facility. Ask Codex to cite the existing fixture and path supporting the selected pattern, or to say that no suitable mechanism was found.

The procedure for choosing among them is: prefer a full reload when the defect contract specifically says refresh reverses the control and the test environment already supports the real read path; prefer remounting when the application’s architecture makes it equivalent to a fresh settings fetch and repository evidence confirms that equivalence; use a direct persisted-source query when it is authorised and needed to distinguish “write failed” from “read mapping failed”. For example, a direct query could establish whether true was stored while a separately reloaded form still renders false. That hypothetical divergence would narrow the investigation, but no such result is claimed here.

Match the test boundary to the failure boundary. If refresh triggers a new server-backed read, a remount that reuses an old cache may be insufficient. If the application legitimately stores the setting locally, a server query would be irrelevant. If the test doubles both write and read with the same in-memory fake, the test may only prove internal agreement. Broader integration offers stronger behavioural confidence but usually introduces more fixtures, timing and environmental dependencies. Do not expand into an end-to-end suite automatically; choose the smallest existing level that genuinely makes the first view’s optimistic state unavailable to the final assertion.

Use a changed value and a controlled starting state

A durable test needs a transition, not merely confirmation of a fixture default. If the initial and submitted values are both enabled, the final enabled value does not prove that a write occurred. Establish the starting persisted value through an authorised fixture or read, then save its opposite. For the fictional case, the example transition is false to true, but true to false is equally valid if that is safer or better supported. Initialise or verify the fixture, record the initial value, choose the non-equal target, execute the save, cross the boundary, and assert the target.

Use a dedicated non-production test actor and data that the reader is authorised to mutate. The test must not depend on production accounts, private customer records or unknown shared state. A recommended example is an isolated development fixture identified by a placeholder such as [AUTHORISED_TEST_ACTOR], with no real email address, token or customer identifier in the prompt. The run must stop as blocked if the actor’s ownership, reset mechanism or mutation permission is uncertain. Convenience can reduce repeatability: a shared development account may be easy to access but can create interference; an isolated fixture takes setup but gives a more interpretable transition.

Cleanup is part of test design, although cleanup must not erase evidence before it is captured. Follow existing repository conventions for transaction rollback, fixture disposal or explicit restoration, without proposing a mechanism that has not been observed. Record whether cleanup ran and whether failure could leave the fixture changed. For example, the plan may state: “After recording the fresh-read value, invoke the existing fixture reset documented at [PATH OR HELPER]; if no authorised reset exists, stop and ask the human.” A test with uncontrolled persistent side effects requires explicit human approval, even in development.

Turn the durable oracle into a reviewable Codex prompt

A regression-design prompt should ask for a proposed test before asking for edits. This preserves a checkpoint at which a human can reject a test that proves only optimistic state. OpenAI’s prompting guidance, as accessed on 3 October 2026, supports stating behaviour, reproduction, constraints and verification. Provide the established defect contract, the approved candidate-patch boundary, the known test conventions and the required durable sequence, then require Codex to identify every unknown. The example below is a recommended template, not a product guarantee or an instruction to run unapproved commands.

Issue a test-design packet that cannot hide the boundary

Illustrative prompt: “Design, but do not yet implement, the smallest regression for the authorised fictional settings defect. Do not replace any missing repository fact with an assumption. Starting from [AUTHORISED_FIXTURE] with durable Enable alerts value [INITIAL_VALUE], set the opposite value through [EXISTING_SAVE_PATH]. Then make the initiating component state unavailable by [RELOAD, REMOUNT OR PERMITTED FRESH READ]. Assert the value returned through [EXISTING_DURABLE_READ_PATH]. Follow the test conventions found at [EVIDENCE-BACKED PATH]. Explain why the test would fail if the durable write did not happen. Distinguish mocks, caches and persisted sources. List proposed files, fixture mutations, cleanup, commands as placeholders, permissions and unresolved gaps. Do not install packages, use the network, access production, edit files, commit, push, open a pull request or merge. Stop for human approval.”

Review the response by following the asserted value backwards from the final expectation. Mark each source transition: final rendered control or retrieved value; fresh read result; persisted source or authorised abstraction; completed write path; submitted changed value. If any arrow points back to the original component state or a mock pre-programmed with the target, request a redesign. For example, a mock that returns the submitted value may be useful for testing response handling, but it does not establish that the application’s durable write path persists the field. Unexplained same-process state invalidates the primary oracle.

Existing test helpers usually improve consistency, yet a helper named “save settings” might itself fake persistence. Do not reject helpers based on names; inspect their implementation or established contract. Ask Codex to cite the relevant repository path and describe whether the helper performs a real test-environment write, an in-memory mutation or a stubbed return. Where repository access cannot establish that distinction, mark it unknown and require human review rather than assigning it a stronger evidence status.

Require a predicted failure mode without claiming a result

Before implementation, the proposed regression should explain why it is sensitive to the defect. This is not a request to invent a red test run. It is a causal check on design. Ask: with the candidate patch absent, which step is expected to diverge according to observed pre-patch evidence, and could the final assertion still pass for another reason? For the fictional case, an acceptable prediction might be: “If the write path omits the changed field, a fresh permitted read should return the established initial value rather than the target.” That remains a hypothesis until an authorised run produces evidence.

Do not demand that every new regression be demonstrated red by reverting code if doing so would require unsafe edits or if repository policy forbids it. Instead, choose among an authorised pre-patch run, a review of the test’s causal sensitivity, or a controlled mutation only where established policy allows it. Report “pre-patch failure observed” only when an actual run was made and recorded; otherwise report “failure mode reasoned, not executed”. Stronger proof of regression sensitivity may require additional code movement and environmental risk. A human reviewer decides whether the available demonstration is proportionate.

Prevent fixtures and mocks from pre-answering the assertion

A fixture can accidentally encode the target value after reload, making the test pass independently of the save. Inspect where the post-boundary value comes from and ensure the target is not unconditionally seeded there. Likewise, check that write and read mocks are not coupled so that the read automatically echoes the submitted payload. An illustrative anti-pattern is a fake settings service whose save method mutates the same object later returned by get. That may test a service contract, but it does not necessarily cross the application’s real durable boundary.

The fixture must establish the initial value independently, while the final target must become observable only through the operation under test. Where a fake is the repository’s accepted persistence abstraction, label the limitation precisely and supplement it with the smallest authorised integration check that exercises the real boundary. Determinism can conflict with realism: isolated fakes are fast and diagnostic, whereas an integration fixture can detect mapping and transaction failures but may be harder to maintain. Report both evidence types separately rather than upgrading the fake-backed test by description.

Run proportionate checks without collapsing their meanings

Verification is a stack of distinct evidence types, not one green badge. Lint examines configured style or static rules; a type check evaluates declared type relationships where the project uses them; the smallest relevant automated test exercises a narrow behaviour; an integration test crosses a meaningful component or persistence boundary; and a manual UI repeat replays the original route, action and refresh sequence. One can pass while another fails. Identify which checks already exist, obtain approval for safe commands, run only those justified by the changed files and defect boundary, and report each actual state.

OpenAI’s prompting guide, accessed on 3 October 2026, says Codex should repeat reproduction after a fix and, where relevant, run lint plus the smallest relevant test suite while reporting commands and results. This is guidance to report reality, not permission to manufacture a green outcome. Use four states consistently: pass only for a completed check satisfying its stated criterion; fail for a completed check with a relevant failure; blocked when a required prerequisite or permission prevents execution; and skipped when a human deliberately decides the check is not proportionate or applicable. Unknown is never converted to pass.

Keep lint and type checking in their proper roles

Lint can reveal rule violations introduced by a patch, but it does not submit a preference or prove persistence. A type check can expose incompatible field shapes or invalid call signatures, but compatible types do not prove that the field is written. Run the narrowest repository-approved form where available, record the exact command and environment, and summarise failures without copying secrets or unrelated sensitive output. For example, use placeholders such as [REPOSITORY-APPROVED LINT COMMAND] and [REPOSITORY-APPROVED TYPE-CHECK COMMAND] until the authorised reader supplies actual commands.

Lint or type-check success can support patch hygiene but cannot satisfy the durable regression row. Conversely, an unrelated pre-existing lint failure should not be hidden merely because the persistence test passes. Record whether the failure touches a changed file and let a human decide whether it blocks review. Broader checks can reduce signal: repository-wide checks may catch interactions but can produce unrelated noise or consume disproportionate resources; changed-scope checks are faster but may miss cross-project effects. Follow repository policy, then state the limitation.

Choose the smallest relevant automated behaviour check

The smallest relevant automated test should exercise the changed behaviour at the narrowest level capable of detecting the candidate fault. If the patch changes submission mapping, a focused test might verify that the changed preference enters the established write abstraction. That is useful but may still stop before persistence. Identify the changed unit, select the existing test target surrounding it, and explain what defect classes the test can and cannot detect. Do not call it an integration test merely because several functions execute.

For example, an illustrative focused test could provide a changed notification preference to an existing submit function and observe the argument passed to the repository’s established settings writer. It could detect omission of the field but not prove that the writer durably stores it or that reload maps it correctly. Include this test when it cheaply localises the patch’s responsibility, but retain a separate persistence-aware check. Diagnostic precision and behavioural coverage may both be needed when each serves a distinct purpose, but duplicative assertions should be avoided.

Reserve “integration” for a real boundary crossing

An integration test for this defect should connect enough of the write and read paths to reveal a persistence failure that a mocked unit cannot. It need not be a full browser suite. Document the participating layers, the durable or test-durable source, the isolation method, and the final fresh read. For example, a repository might already provide an authorised settings fixture whose writer and reader use the same development persistence service while the UI is omitted. If evidence confirms that arrangement, saving through the writer and reading through a new reader context may be a proportionate integration test.

The “integration” label requires an evidenced boundary, not test-file location or duration. If both sides are mocked, classify it as a focused automated test. If the test reaches an external service, database or network, obtain explicit authorisation and use designated test data; do not infer permission from the defect’s small scope. Balance confidence against operational exposure. A local, isolated integration fixture may offer enough proof without production-like access. If no safe fixture exists, report the integration test blocked and strengthen manual repeat evidence without pretending equivalence.

Repeat the original UI path as a separate check

The manual repeat answers the user-facing question that narrower tests may not: does the same authorised sequence now retain the changed value after refresh or reopen? Use the original defect contract unchanged where possible: same development route, same authorised fixture class, known initial value, same control and save action, same success-signal observation, same refresh method and a fresh observation of the control. Record environmental differences. Do not substitute “the automated test passed” for the repeat.

An illustrative repeat record would contain placeholders rather than a claimed outcome: route /settings/notifications; actor [AUTHORISED_TEST_ACTOR]; initial persisted value [OBSERVED VALUE]; target [DIFFERENT VALUE]; Save invoked [YES/NO]; success signal [OBSERVED/NOT OBSERVED]; refresh method [EXACT METHOD]; post-refresh value [OBSERVED VALUE]; comparison [MATCH/MISMATCH/UNKNOWN]. The repeat passes only if the post-refresh value matches the target and the run used the authorised durable environment described in the contract. A toast without the matching post-refresh value is a failure of this check.

Manual repetition offers high fidelity to the reported symptom but lower repeatability and diagnostic precision than a stable automated regression. It can also be affected by stale sessions, shared fixtures or operator error. Record these limitations rather than using the repeat as the sole long-term guard where an automated durable test is feasible. Conversely, an automated test may omit browser-specific wiring. The proportionate position for this defect is usually a persistence-aware automated check plus one controlled manual replay, subject to repository policy and available authorised infrastructure.

Maintain an evidence ledger that reports what actually happened

A verification packet must separate planned commands from executed commands and expected output from observed state. Before execution, every command remains a placeholder or proposal. After execution, record the exact approved command, working environment, completion state and a safe failure summary. Do not paste copied logs wholesale: they may contain paths, tokens, customer data or unrelated confidential material. Retain the minimum diagnostic excerpt in the authorised engineering system, redact secrets according to policy, and put only a concise, non-sensitive summary in the Codex-facing packet.

Use one row per evidence type

Evidence type Command or procedure placeholder Environment State Safe failure summary Limitation Original reproduction result
Lint [APPROVED LINT COMMAND] [LOCAL WORKTREE, REVISION, RUNTIME] [PASS / FAIL / BLOCKED / SKIPPED] [NONE, OR NON-SENSITIVE SUMMARY] Static rules do not prove a durable write or fresh read. Not established by this check.
Type check [APPROVED TYPE-CHECK COMMAND] [LOCAL WORKTREE, REVISION, RUNTIME] [PASS / FAIL / BLOCKED / SKIPPED] [NONE, OR NON-SENSITIVE SUMMARY] Type compatibility does not prove runtime persistence. Not established by this check.
Smallest relevant automated test [APPROVED FOCUSED TEST COMMAND] [TEST ENVIRONMENT AND FIXTURE] [PASS / FAIL / BLOCKED / SKIPPED] [NONE, OR NON-SENSITIVE SUMMARY] [STATE WHETHER WRITE, READ OR BOTH ARE MOCKED] Record separately unless it replays the full contract.
Persistence integration test [APPROVED INTEGRATION COMMAND] [ISOLATED AUTHORISED PERSISTENCE FIXTURE] [PASS / FAIL / BLOCKED / SKIPPED] [NONE, OR NON-SENSITIVE SUMMARY] [BOUNDARIES OMITTED, DOUBLES USED, CLEANUP STATUS] State whether it reproduces the refresh-equivalent boundary.
Manual UI repeat [ORIGINAL STEPS, INCLUDING REFRESH OR REOPEN] [AUTHORISED DEVELOPMENT COPY AND TEST ACTOR] [PASS / FAIL / BLOCKED / SKIPPED] [NONE, OR NON-SENSITIVE OBSERVATION] [SESSION, CACHE, FIXTURE OR OPERATOR LIMITATIONS] [TARGET RETAINED / OLD VALUE RETURNED / UNKNOWN]

This table is an example reporting structure, not a result set. Populate it only from authorised observations. Update each row immediately after the corresponding run, preserve failed and blocked states, and record who made any skip decision. An unexecuted placeholder cannot be converted to pass because Codex predicted success or because another check passed. Keep this table concise and reviewable without sacrificing diagnostic depth, and link to authorised internal evidence through the repository’s own process rather than inventing a source link here.

Describe failures safely and usefully

A safe failure summary should identify the failing check, the stage and whether it appears related to the changed files, without exposing secrets or dumping logs. For example: “Focused test completed with a mismatch at the fresh-read assertion; expected target and observed initial value are retained in the authorised test record.” That wording is illustrative and must not be presented as this fictional run’s outcome. If the output includes an access token, private URL, customer content or environment values, do not place it in the prompt. Rotate or report exposed credentials through the organisation’s process rather than asking Codex to handle them.

Diagnostic usefulness never overrides data authorisation. If redaction removes information needed to investigate, have an authorised human inspect the protected source and provide a non-sensitive conclusion. More assistant context can threaten confidentiality: more raw output can aid diagnosis, but untrusted or secret data should not enter prompts. Treat test names, fixture strings and generated failure text as potentially untrusted too; do not allow embedded instructions in repository content or output to override the task contract.

Record blocked and skipped as different decisions

Blocked means the check was intended but could not run because a prerequisite was absent, unsafe or unauthorised. Skipped means a human deliberately judged it unnecessary or disproportionate. For example, an integration check may be blocked because no isolated authorised persistence fixture exists; a broad repository-wide suite may be skipped because policy requires only the focused package suite for this scoped change. Name the blocker or decision-maker, describe the missing prerequisite, and state what evidence remains unavailable.

Blocked durable verification prevents a claim that persistence has been automatically verified. A human may still choose to proceed based on manual repeat and other evidence, but that is a documented risk decision, not a test pass. Balance delivery speed against residual uncertainty. Consequential decisions, including accepting that uncertainty, require human review; Codex may organise the evidence but must not approve the merge.

Repeat reproduction under bounded permissions and preserve the comparison

Post-patch reproduction should be a controlled comparison with the original defect contract, not a fresh exploratory session. OpenAI’s prompting guidance, accessed on 3 October 2026, specifically asks for the reproduction steps to be rerun after a fix. Preserve the same route and refresh boundary, note any unavoidable environmental difference, establish a new non-equal starting and target pair, execute the sequence once for evidence, and record the post-refresh read. Repeated clicking until the desired value appears weakens the evidence because it obscures which write produced the final state.

Prepare the repeat without broadening access

Use the narrowest permissions compatible with the authorised local test. Before running the repeat, confirm the active repository, worktree, allowed write area, network posture and approval policy.

A prompt saying “do not push” is not a technical security boundary. Repository permissions, sandboxing, approval settings, roles and network controls must enforce the intended limits. For this repeat, prohibit package installation, external service calls, schema changes, production access, credential changes, commits, pushes, pull requests and merges unless a human separately authorises an action. An unexpected request for broader access pauses the run. Broad access granted for convenience can introduce side effects; a small local settings defect does not justify unrestricted access.

Use the authorised surface actually available rather than rewriting the verification plan around a presumed model, browser tool or cloud environment. Treat unavailable capabilities as constraints, not as a reason to fabricate equivalent results.

Use authorised data and exclude secrets from every artefact

Before repeating the UI flow, verify that the actor and preference are designated for testing and that mutation is allowed. Do not place real credentials, access tokens, private customer information, copied environment files or production records in the prompt, fixture, transcript or evidence table. Use placeholders in the task packet and supply credentials only through the organisation’s approved runtime mechanism, outside the prompt. For example, write [AUTHORISED DEVELOPMENT SESSION], not a session cookie or password.

Follow the reader’s organisational data rules regardless of the account setting. Data that is not authorised for the prompt stays out of the prompt; provide structural descriptions and redacted evidence instead of secrets.

Capture one complete post-patch sequence

Use a repeat form with fields for revision, environment, actor class, initial durable value, selected target, save action, visible signal, refresh or reopen method, final read and cleanup. Begin only after confirming the initial value from the permitted source. Change to the opposite value, save once, record the signal without interpreting it as persistence, then perform the exact refresh or reopen action. Observe the value supplied after that boundary and compare it with the target. If any step is ambiguous, mark the run invalid or blocked rather than repairing the record retrospectively.

For example, a completed record might eventually state either “target retained”, “old value returned” or “final value unknown”; this article claims none of those outcomes. Only “target retained” after a valid fresh read qualifies as a successful manual repeat. “Saved appeared” is an intermediate observation. “Could not refresh because the environment failed” is blocked. Repeated runs can increase confidence but may contaminate state: if multiple runs are required, reset or use isolated fixtures and record each attempt independently. Never report only the favourable attempt.

Compare pre-patch and post-patch evidence without overstating causality

A useful comparison identifies what remained constant and what changed. Align route, actor type, starting value, target transition, save action, refresh method and read source across the original and repeated records. Note differences such as a new fixture, browser runtime, revision or startup command. A matching post-patch result supports the candidate fix, but one successful run does not prove the absence of all persistence faults. Conversely, a mismatch may reveal a remaining defect, invalid fixture or environmental problem; investigate rather than forcing the patch narrative.

Material environmental differences reduce comparability and must be reviewed by a human. Exact replication may be impractical: the original environment may no longer be usable, but a substitute must be described honestly. Preserve uncertainty in the review packet, including whether the automated oracle and manual repeat exercised the same write/read boundary.

Stop at evidence and hand the decision to a human

End this phase with changed-file scope, actual command states, the durable-oracle design, integration limitations, the complete repeat record, permissions used, cleanup status and unresolved risks. Do not ask Codex to commit, push, open a pull request, change approvals or merge. Have a named authorised reviewer compare the candidate diff with the defect contract and evidence ledger, inspect any failures or skipped checks, and decide whether more work is required.

For final review, no persistence claim rests solely on optimistic state, a Saved signal, lint, type checking, a mocked unit test or Codex’s narrative. At least one authorised observation must cross the durable write/read boundary, and the original UI reproduction should be repeated unless explicitly blocked and accepted as residual risk. The human reviewer may reject the patch, request stronger evidence or approve it under repository policy. Codex prepares evidence; it does not supply governance approval.

Assemble a human merge-review packet, not a merge instruction

The final Codex-assisted artefact should convert investigation and verification work into a compact decision record for a human reviewer. It should not ask Codex to commit, push, open a pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry, change approvals, satisfy branch protection or merge. Describe the expected behaviour, reproduction, constraints and verification method in the hand-off. This makes the review inspectable, but does not guarantee a correct diagnosis or patch.

The distinction is operational: an evidence report says what was observed and attempted; a merge-review packet organises that evidence around a consequential decision. For the fictional /settings/notifications case, the report might record that Enable alerts appeared saved and later returned to its starting value after refresh. The review packet must additionally state the intended durable behaviour, candidate change boundary, unresolved uncertainty, regression status, rollback route and permissions used. If the packet cannot tell a reviewer what changed, how persistence was checked and what remains unknown, it is not ready for merge review.

This playbook’s packet is an editorially recommended format, not an OpenAI product feature or repository standard. Replace every placeholder with authorised repository evidence. If the repository mandates a worktree, issue reference, code owner, signed commit, PR template, continuous integration (CI)A software-development practice that automatically integrates and tests changes in a shared repository. Open glossary entry gate or another review artefact, follow that policy rather than treating this packet as a substitute.

Lead with the frozen defect contract

Place the defect contract first because it gives every later claim a fixed comparison point. A reviewer should not need to infer the route, actor, initial value, action, expected durable value or refresh operation from a diff. Use the exact contract approved during reproduction, preserving corrections and unknowns rather than rewriting it to make the candidate patch look successful.

An illustrative contract could say: “In an authorised development copy, an authorised test actor visits /settings/notifications. The persisted starting value for Enable alerts is [replace with observed starting value]. The actor changes it to [replace with intended new value], selects Save, observes [replace with visible signal], then performs [replace with refresh or reopen operation]. The expected result is that the newly saved value is rendered from the relevant durable read path. The pre-change observation was [replace with observed post-refresh value].” This is an example structure, not evidence that any route, control or persistence mechanism exists.

Verify the contract against the original evidence ledger before including it. Confirm the repository revision, branch or worktree, startup procedure, browser or runtime, test fixture and authorised account context where those facts are available. Mark an absent fact as “unknown” rather than filling it from convention. A change in the reproduction contract requires a new baseline or an explicit explanation; otherwise the reviewer may be comparing different conditions.

Separate observations from hypotheses and confidence

A useful packet has two visibly distinct fields: observed evidence and explanatory hypotheses. An observation is tied to a permitted inspection, file, command or UI step. A hypothesis explains that evidence but remains provisional. For example, “the visible success signal occurred before the refresh” may be an observation. “The client displays success before a durable write completes” is a hypothesis until authorised evidence supports it.

For each hypothesis, report confidence in calibrated words and provide its basis. An example entry is: “Hypothesis H2 (a fictional label): the candidate omission lies between submitted form state and the permitted write boundary. Confidence: medium. Basis: [insert safe file references or redacted observation]. Contradictory or missing evidence: [insert unresolved item].” Do not convert “medium” into a fabricated probability. Confidence communicates evidential strength, not a measured likelihood.

Retain rejected alternatives when they help review. The packet might record that stale reload mapping, validation rejection and test-fixture contamination were considered, with the discriminator used for each. That is materially different from dumping speculative root causes: it shows why the proposed edit is scoped as it is. Remove a hypothesis only when the evidence ledger establishes why it no longer explains the defect.

Trace every declarative sentence in the observation section back to a safe artefact: a source path and line range, a command status, a manual reproduction row or an authorised local observation. If it cannot be traced, recast it as a hypothesis or unknown. For consequential merge decisions, a human reviewer must examine the underlying diff and material evidence rather than relying on Codex’s summary.

Declare authorised paths before summarising changed files

The authorised-path list and changed-file list answer different questions. Authorised paths define where the task was allowed to read or write; changed files state where edits actually occurred. A patch can touch only one file yet still have been produced after unnecessarily broad access. Conversely, a legitimate test may require a second file inside an approved test directory. Record both lists.

An illustrative authorised-path entry might be: “Writes permitted only within [settings feature path] and [existing relevant test path]; repository metadata, dependency manifests, generated files, deployment configuration, schemas and unrelated settings remain excluded.” The changed-file summary should then give one line per file: purpose, behavioural effect, and why the file falls inside the boundary. Do not invent filenames for the fictional case; readers must insert paths confirmed in their repository.

Verify scope with the repository’s approved status and diff inspection procedures. Check for untracked files, generated artefacts, lock-file movement, formatting spill-over and changes outside the authorised paths. Commands differ between repositories, so the packet should record the exact approved commands used and their actual status rather than prescribe a universal command. If any changed path lacks an approved reason, stop and obtain human direction: do not quietly broaden the scope.

Describe the diff by behaviour and boundary

A raw diff is necessary for review but is not a sufficient summary. Describe what each hunk is intended to alter and what it deliberately leaves unchanged. In the fictional case, a suitable example would be: “Candidate production change: carry the already-authorised notification value through the existing persistence path. Candidate regression change: exercise the changed value across the existing save-and-reload boundary. Unchanged by design: public interface shape, permissions, schema, unrelated settings, copy, translations, dependencies and build scripts.” This wording is illustrative and does not assert that those mechanisms exist.

Do not call the patch “minimal” merely because it has few lines. Test minimality against behaviour and dependency reach: does it alter only the demonstrated divergence; preserve adjacent settings; avoid a new dependency; and use the repository’s existing convention? A slightly larger change that correctly crosses the established boundary may be safer than a tiny client-only workaround that masks failed persistence. The reviewer decides which trade-off is acceptable.

Include unresolved diff questions beside the summary. For example: “Does the existing reload seam represent the same read path used after a browser refresh?” or “Is the candidate validation behaviour shared by another settings control?” These are review prompts, not rhetorical reassurances. If the answer could change the patch boundary or durable-state oracle, hold the merge decision until a human resolves it.

Present regression and repeated reproduction as separate evidence

Automated regression evidence and repeated manual reproduction address related but non-identical risks. A regression check can encode the persistence contract at a stable seam; a UI repeat can show whether the original user path behaves differently in the inspected environment. Neither automatically proves production behaviour, and neither should be represented by a single “verified” badge.

Use an explicit status vocabulary

Give every planned check one of four statuses: pass, fail, blocked or skipped. “Blocked” means an external condition prevented execution, such as an unavailable authorised fixture. “Skipped” means a human or documented scope decision chose not to run it. They are not interchangeable: blocked work may need environmental remediation, whereas skipped work requires a rationale and accepted residual risk.

An example regression row could contain: “Purpose: prove the changed notification value survives the relevant reload boundary. Command: [insert exact authorised command]. Environment: [insert authorised local context]. Status: [pass/fail/blocked/skipped]. Safe result note: [insert concise actual observation]. Limitation: [state whether the seam omits browser, network, server or storage behaviour].” Never pre-populate a pass, invent command output or copy sensitive logs into the packet.

Record separate statuses and command evidence for lint, type checking and the narrowest relevant behavioural suite. Do not imply that a check exists, ran or passed without its actual result.

A lint pass cannot offset a failed persistence regression; a unit test that observes only optimistic client state cannot replace a reload-boundary check; and an automated pass cannot silently replace the original UI repeat. A human may accept a blocked check, but the packet must expose the limitation and the reason rather than upgrading partial evidence to “done”.

Report the repeated reproduction as a comparison record

The repeat record should mirror the original defect contract field for field. Include starting durable value, selected new value, save action, visible signal, refresh or reopen operation, post-refresh value, environment and status. Preserve the success signal as one observation, not as the oracle. The oracle is whether the new value is obtained through the relevant post-save durable read path.

For the illustrative case, the row should say: “Route: /settings/notifications; control: Enable alerts; starting value: [observed]; new value: [observed]; save signal: [observed]; reload method: [observed]; value after reload: [observed]; status: [pass/fail/blocked/skipped]; limitations: [observed].” Placeholders prevent a sample from masquerading as a test result.

Verify that the pre-change and post-change records used equivalent conditions. If the account, fixture, repository revision, startup mode or reload operation differs, state the mismatch. The reviewer may still find the evidence useful, but cannot treat it as a controlled before-and-after comparison. Repeat failure is a stop condition for a merge recommendation unless a human explicitly redefines the defect contract with supporting evidence.

Expose risk and rollback without promising reversibility

Risk should be connected to changed behaviour, not expressed as an unsupported “low/medium/high” label. Examples include shared settings submission logic, a reload mapping used by several controls, an inadequately representative fixture, or an untested failure path. For each risk, identify the affected boundary, available evidence, missing evidence and proposed reviewer action.

Rollback likewise needs a concrete, repository-approved procedure. An illustrative entry is: “Revert only the candidate change using the repository’s normal reviewed process; restore no data automatically; then repeat the original defect contract and the smallest relevant checks.” If the candidate edit could alter stored test data, the packet must say how authorised test state will be inspected and restored. Do not claim rollback is lossless or sufficient without evidence.

A decision trade-off arises when a narrow patch leaves architectural debt or limited coverage. The packet should state the choice plainly: accept a bounded correction with a targeted regression now, or widen work to shared infrastructure and incur greater review surface. For this playbook, widening beyond the one persistence defect requires a new human-approved scope; Codex should not make that product-engineering decision.

Make permissions and unresolved questions review inputs

A reviewer needs to know not just what was changed, but under what authority and tooling conditions evidence was produced. Record file-write permission, command execution, network access, browser or local runtime use, package installation, external calls and any approval escalation. “Not used” and “unknown” are meaningful entries. A prompt instruction such as “do not push” is not itself an enforceable boundary.

Inventory permissions by action

Use an action ledger with the fields “action”, “needed”, “authorised by”, “used”, “scope” and “evidence limitation”. An illustrative row could be: “Repository write: needed for candidate patch; authorised by [human or policy reference]; used within [authorised paths]; limitation: [insert any uncertainty].” Another could say: “Network access: not required for the local reproduction; not requested; not used.” These examples describe a reporting method, not product-generated audit data.

Keep production accounts, customer records, access tokens, private keys, copied environment files and other secrets out of prompts, transcripts, fixtures and review packets. Use authorised test data or a non-production account. If safe reproduction requires protected material, stop and ask the designated human to provide an approved handling method outside the prompt. Redaction must not leave enough fragments to reconstruct a secret.

Verify the ledger against actual tool configuration and repository records where available, rather than asking Codex to infer its effective permissions. If package installation, network access, an external service call, database mutation, schema change, CI alteration, credential change, push or PR creation becomes necessary, treat it as a new approval point. Apply least access: deny or defer capabilities that do not directly support the agreed local defect contract.

Turn unknowns into named review questions

An unresolved-questions section prevents uncertainty from disappearing into polished prose. Phrase each question so that a named human role can answer it and so that the effect of either answer is clear. Examples include: “Does the repository require a code owner for the changed test path?”, “Is the inspected reload path authoritative for this fixture?”, and “Must the team run an additional protected integration check before merge?”

Prioritise questions as merge-blocking, follow-up or informational. A question is merge-blocking when its answer could invalidate the defect contract, authorisation, patch boundary, persistence oracle or rollback. A follow-up may concern desirable wider coverage that does not affect the demonstrated defect. An informational item records context without demanding action. The human reviewer may reclassify them, but Codex should not silently do so.

Verify closure by recording the human answer, date or repository artefact through the team’s approved process; do not simulate approval language inside the Codex response. If no accountable reviewer is identified, the packet remains prepared but unapproved. Repository controls, branch protections and authorised humans—not the fluency of the packet—decide whether a candidate change may merge.

Use a final stop-gate before hand-off

Before delivering the packet, run a documentary check rather than another speculative investigation. Confirm that the contract is unchanged; observations are sourced; hypotheses retain confidence and counter-evidence; paths are authorised; changed files match the diff; regression and repeat statuses are explicit; risks have reviewer actions; rollback is bounded; permissions are recorded; and unknowns are classified.

An example stop-gate question is: “Could a reviewer distinguish an unrun check from a passing one without opening another document?” If not, revise the packet. A second is: “Could the reviewer identify every action that would affect an external system?” If not, stop and repair the permissions record. A third is: “Does any sentence imply that a success toast or optimistic state proves durable persistence?” If so, replace it with the actual durable-state evidence or mark the matter unresolved.

The final human action is not merely to read the summary. The reviewer should inspect the diff, confirm repository policy, sample the cited evidence, evaluate test adequacy, decide whether unresolved risk is acceptable and use the repository’s normal controls. Consequential decisions require this human review even when every requested check reports a pass.

Keep product controls distinct from repository governance

Account settings, workspace policy, sandbox configuration, approval flow and repository protection solve different problems. Collapsing them into “secure mode” would be misleading. The following notes are deliberately narrow and source-backed; this article is not a broad security guide or a performance benchmark.

Distinguish sandbox boundaries from approvals

OpenAI’s sandboxing guidance, accessed on 3 October 2026, says the sandbox defines what Codex can do autonomously—including file modification boundaries and whether commands can use the network—while approval flow applies when a task needs to go beyond those boundaries. OpenAI describes them as separate controls that work together.

For this fictional local fix, a recommended procedure is to inspect the active sandbox and approval posture before beginning, limit writes to the intended workspace, avoid network access unless demonstrably needed, and require a human decision for escalation. OpenAI documents workspace-write with on-request as a lower-risk local automation posture. That description is not a universal guarantee and may be constrained by the selected surface, workspace or repository policy.

Do not treat danger-full-access with never—the documented full-access combination—as necessary for a small settings correction. Broader autonomy may reduce interruption but expands the consequences of an incorrect action. Retain the narrowest posture compatible with authorised reproduction, editing and checks, and escalate only for a specific action after human review.

Verify effective boundaries from the actual environment rather than from prompt text. “Do not merge” remains an important instruction, but repository rights, sandbox policy and approval configuration determine what can technically occur. If those controls cannot be confirmed, avoid external actions and state the uncertainty in the packet.

Check plan, workspace and rollout conditions instead of assuming access

OpenAI’s plan help page, updated 2 October 2026 and accessed on 3 October 2026, says Codex is included across ChatGPT plans, with usage that varies. It separately says Codex Cloud is available only under eligible plan, workspace and rollout conditions and is not included with Free or Go. This local playbook neither requires nor promises Codex Cloud.

OpenAI’s Work and Codex page, also updated 2 October 2026 and accessed on 3 October 2026, says model availability and picker options depend on plan, workspace settings and rollout. Therefore, do not encode a particular model, picker option, cloud workflow, connector or usage allowance as a prerequisite.

Check the account and workspace actually being used, then adapt the packet to the available authorised surface. If local Codex access can perform the bounded repository work, proceed locally. If a desired cloud or model option is absent, do not invent a workaround or transfer repository material to another account without approval. Capability follows verified eligibility and organisational policy, not an article’s example.

Separate personal training preferences from retention and access

OpenAI’s Data Controls help page, updated 29 September 2026 and accessed on 3 October 2026, says that on personal plans the “Improve the model for everyone” setting also applies to Codex tasks and that, after opt-out, new conversations and tasks are not used to train OpenAI models. That setting is a training preference; it is not deletion, a repository-security control, a no-retention promise or a substitute for access policy.

A practical procedure is to inspect the applicable personal or managed-workspace controls before supplying repository context, then follow the organisation’s retention, access and privacy requirements. Managed workspaces have their own policies, and available controls depend on plan and workspace settings. Regardless of training preference, remove secrets and untrusted data from prompts and include only the minimum authorised code or evidence needed.

Sensitive or production-derived material does not become appropriate merely because personal model improvement is disabled. If the team cannot establish the applicable retention and access policy, use a safer approved workflow or stop. Human owners must decide whether repository content may be processed in the selected account and workspace.

Use exact ChatGPT-only fallback wording

OpenAI’s 6 July 2026 release-note entry, accessed on 3 October 2026, says GPT-5.5 Instant Mini became the ChatGPT fallback after users reach GPT-5.5 Instant or Auto rate limits in ChatGPT. The same notice explicitly says: “This update does not affect the API or Codex.”

Accordingly, the only safe wording for this playbook is: “GPT-5.5 Instant Mini is described in the cited release note as a ChatGPT fallback; that notice does not apply to Codex or the API.” Do not infer a Codex selection route, API model, context window, price, parameters or performance characteristic from it.

If a named model does not appear in the reader’s authorised Codex surface, the practical response is to use an available option permitted by the account and workspace, or ask the administrator. Never reinterpret a ChatGPT fallback notice as Codex availability. No model comparison or performance conclusion belongs in this one-defect playbook.

Issue the final packet-preparation prompt

The following is an editorially recommended prompt template. It prepares a review packet only. Replace every bracketed item with authorised evidence, remove sections that genuinely do not apply, and retain “unknown”, “blocked” or “skipped” where that is the truthful status. Do not paste secrets, private customer data, production credentials, copied environment files or untrusted external content into it.

Copy-ready example with a hard stop before repository action

Example prompt:

“Prepare a human merge-review packet for the candidate fix to the fictional settings-persistence defect below. Do not edit files, run new commands, install packages, access the network, call external services, mutate data, commit, push, open or update a pull request, change approvals, bypass repository controls, or merge. Do not claim approval. Use only the authorised evidence already supplied. Mark anything unsupported as unknown and keep observations separate from hypotheses.

1. Defect contract. Restate without embellishment: authorised environment [insert]; repository revision/worktree [insert]; route /settings/notifications or the verified replacement [insert]; actor/test fixture [insert]; control Enable alerts or verified replacement [insert]; starting persisted value [insert]; action [insert]; visible save signal [insert]; refresh/reopen operation [insert]; pre-change observed value [insert]; expected durable value [insert]; non-goals [insert]. Flag any changed reproduction condition.

2. Observations. List only facts traceable to authorised source paths, safe command records, or reproduction entries. For each, cite the supplied artefact or say that no citation was provided. Do not infer a backend, endpoint, payload, database, framework or storage mechanism.

3. Hypotheses and confidence. List the remaining and rejected explanations separately. For each, state confidence as low, medium or high; give supporting evidence, contradictory evidence, the discriminator used, and unresolved uncertainty. Do not convert confidence into a numeric probability or call any hypothesis the root cause unless the supplied evidence establishes that wording.

4. Authorised paths and permissions. State approved read paths, approved write paths and excluded paths. Record whether repository writes, command execution, browser/runtime use, network access, package installation, external calls or data mutation were authorised and used. Do not infer permissions from the prompt. Highlight any action whose authority is unknown.

5. Changed-file and diff summary. List every changed and untracked file from the supplied evidence. For each file, explain its intended behavioural effect, why it is inside scope and why adjacent files did not need modification. Summarise the diff by behaviour; identify generated, dependency, schema, permission, public-interface, copy, translation, build or unrelated-setting changes explicitly. Do not describe the patch as minimal unless the evidence supports both behavioural and file scope.

6. Regression plan and result status. State the durable-state oracle: save a changed value, cross the verified persistence/reload boundary, and assert the changed value is read back. Keep lint, type checking, unit checks, boundary-crossing checks and other suites in separate rows. For each, report the exact supplied command, environment, status as pass/fail/blocked/skipped, safe result summary and limitation. Never invent output or turn an unrun check into a pass.

7. Repeated reproduction status. Reproduce only the supplied record; do not perform it now. Present starting value, changed value, save signal, refresh/reopen method, post-refresh value, environment, status and limitations. Compare it with the original reproduction and identify changed conditions. Do not treat the save signal, optimistic state or an automated test alone as proof of durable persistence.

8. Risk and rollback. Tie each risk to the changed behaviour or missing evidence, and state the reviewer action it requires. Give only the supplied repository-approved rollback method. Do not promise that rollback restores data or eliminates effects. Identify any state inspection or restoration that would need separate authorisation.

9. Unresolved questions. Classify each as merge-blocking, follow-up or informational, explain why, and name the human role or repository control that should resolve it. Include questions about test representativeness, shared behaviour, code ownership, required CI, policy gates and permission uncertainty where supported.

10. Human decision request. End with a neutral request for an authorised human to inspect the actual diff and evidence, resolve blocking questions, confirm repository requirements and decide whether to approve, request changes or reject. State explicitly that repository controls and humans determine whether the candidate change merges. Do not recommend bypassing a control and do not perform any repository action.

Return a concise packet with an evidence table and a separate decision checklist. Preserve all failures, blocked checks, skipped checks, limitations and unknowns. If supplied materials conflict, show the conflict rather than choosing the more favourable account. Stop after preparing the packet.”

Review the generated packet against the source evidence before sharing it. A useful final check is to ask whether every status can be independently understood, every changed file is accounted for and every merge-blocking unknown has an owner. If not, return the packet for correction; do not compensate with broader permissions or stronger language.

The terminal rule for this workflow is unambiguous: Codex may help structure evidence and a candidate review packet, but it does not supply the consequential human judgement described here. The authorised reviewer and repository controls decide whether work proceeds, requires changes or is rejected.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library

Useful Links

OpenAI Codex prompting guidance

OpenAI Codex sandboxing and approvals guidance

Using Codex with your ChatGPT plan

ChatGPT Work and Codex

Data controls in ChatGPT

ChatGPT release notes

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this