Safer Codex App Server Clients: Transport, Events and Recovery

Conceptual dark workstation connected through a locked gateway to a server node by separate pipe, socket and encrypted paths.

Start With the Smallest Reachable Trust Boundary

This guide considers a constrained internal tool: a developer-operated client that runs for a known purpose, on a controlled workstation, with deliberately limited reach. It is not a general service, a public endpoint or a continuous integration (CI)A software-development practice that automatically integrates and tests changes in a shared repository. Open glossary entry runner. Its first design question is therefore not which transport appears most flexible, but which transport creates the smallest boundary that still permits the required work.

The official Codex documentation draws a useful product boundary: “Use the Codex software development kit (SDK)A collection of libraries, tools and documentation for building against a platform. Open glossary entry to automate coding tasks, including jobs in CI.” It directs custom clients that manage authentication, conversation history, approvals and streamed agent events to the Codex App Server instead. As documented by OpenAI and reviewed on 6 October 2026, the App Server is the stateful client surface addressed here. The Codex SDK is the separate choice for CI-style coding jobs. This guide does not turn the App Server into a CI mechanism or compare every available Codex integration.

Conceptual dark workstation connected through a locked gateway to a server node by separate pipe, socket and encrypted paths.
An original conceptual view of choosing a constrained client transport boundary.

Evidence checkpoints

Documented point: Choose the Codex app server when a custom client must manage authentication, conversation history, approvals and streamed agent events, while directing CI-style coding jobs to the Codex SDK. This guide addresses the documented custom-client surface rather than CI job automation. [OpenAI documentation: Codex SDK]

Documented point: For a same-host child process, the default stdio transport uses newline-delimited JavaScript Object Notation (JSON)A text format for representing structured data as objects, arrays, numbers, strings and other values. Open glossary entry, whereas WebSocket transport is experimental and unsupported and should be selected only after its exposure is deliberately controlled. An experimental and unsupported status does not establish a production-support guarantee. [OpenAI documentation: Codex App Server]

Documented point: On each transport connection, send one initialize request before any other method and then send the initialized notification, because earlier requests receive a Not initialized error. [OpenAI documentation: Codex App Server]

Documented point: Keep thread, turn and item state separate: a thread contains turns, a turn contains items and incremental updates, and a retained thread identifier supports a deliberate resume design after disconnection. Resumption design is a client procedure, not an exactly-once delivery guarantee. [OpenAI documentation: Codex App Server]

Documented point: Render item deltas as provisional progress but persist the completed item as final state, because a final plan item may differ from concatenated deltas. [OpenAI documentation: Codex App Server]

Documented point: Treat an approval as a server-initiated JavaScript Object Notation Remote Procedure Call (JSON-RPC)A stateless, lightweight remote procedure call (RPC) protocol that uses JavaScript Object Notation (JSON) as its data format. Open glossary entry request and show its command and scope to a human decision-maker, using the supplied thread and turn identifiers to scope the interface state before returning a decision. Approval decisions must remain meaningful human consent rather than a silent auto-accept rule. [OpenAI documentation: Codex App Server]

Documented point: For a non-local WebSocket connection, require Transport Layer Security (TLS)A protocol that lets client/server applications communicate over the Internet in a way designed to prevent eavesdropping, tampering and message forgery. Open glossary entry and WebSocket authentication, present a bearer credential during the handshake and avoid raw bearer tokens in process arguments, URLs, logs and repositories. Non-loopback WebSocket listeners currently allow unauthenticated connections by default during rollout. [OpenAI documentation: Codex App Server]

Documented point: When the server returns JSON-RPC error -32001 with Server overloaded; retry later., retry only with capped exponential delay and jitter, and evaluate the network sandbox separately from approval timing because they are different controls. Backoff guidance does not promise exactly-once delivery after a disconnect. [OpenAI documentation: Codex App Server OpenAI documentation: Agent approvals and security]

Terms in the related guides: A command-line interface (CLI)A text-based interface for running commands and tools. Open glossary entry is a text-command entry point; a user interface (UI)The controls and visual surfaces through which a person interacts with software. Open glossary entry is the operator-facing display. The two views can report different evidence about an operation, so the approval record must say which was observed.

A narrow client still carries consequential responsibilities. It will eventually need to initialise a connection, associate events with the right conversation, present proposed commands and scopes for human review, and recover without inventing completion state. Those later responsibilities do not justify making the transport broadly reachable. Reachability should follow an identified requirement, not anticipated convenience.

The working example throughout this guide is fictional. “Harbour Desk” is a suggested design for a developer-only review helper on one workstation. It starts a Codex App Server for one user session, displays streamed review activity and will later present approval requests to that developer. It has no browser-facing service, no second workstation to support and no requirement for unattended CI execution. Those constraints make its transport decision concrete rather than hypothetical.

A useful initial procedure is to write five facts before implementing any protocol methods:

  1. Process ownership: record whether the client starts and stops the server process, or connects to a process managed elsewhere.
  2. Same-host requirement: state whether every legitimate client and server process runs on the same operating-system host.
  3. Exposure: identify which users, processes, interfaces and networks could reach the selected channel.
  4. Credential handling: determine whether the channel needs a credential and, if so, which secret-management boundary supplies it without exposing it in code, a Uniform Resource Locator (URL)The address used to identify and access a resource on the web. Open glossary entry, a command-line argument, a log, a screenshot or an example repository.
  5. Testability: specify how a reviewer will prove both a permitted connection and the absence of an unintended non-local route.

For Harbour Desk, the suggested answers are: the desktop client owns the child process; both ends remain on one workstation; only the parent and child need to exchange protocol messages; no network credential is needed for the child’s standard streams; and the test harness can inspect process ancestry, open listeners and shutdown behaviour. These are example design inputs, not claims about a product guarantee.

Choose the least reachable transport that satisfies the recorded same-host and process-ownership requirements. A spawned child connected through standard input and output is preferable when the client owns the process. A deliberately local inter-process communication boundary may be justified when separately managed same-host processes must communicate. A network transport is justified only when a real network boundary must be crossed. Ease of attaching another client later is not, by itself, sufficient reason to expose a listener now.

Map the client-to-server path before choosing a transport

A transport diagram and a protocol diagram answer different questions. The transport diagram shows who can establish a channel; the protocol diagram shows what messages flow after that channel exists. Treating JSON-RPC initialisation as though it controlled network admission would confuse these layers. A peer that can reach an inadequately protected listener has already crossed the transport boundary before any JSON-RPC method is considered.

OpenAI’s App Server documentation, as reviewed on 6 October 2026, describes standard input/output as the default transport: stdio carries newline-delimited JavaScript Object Notation, commonly called JSON Lines (JSONL)A text format containing one valid JavaScript Object Notation value per line. Open glossary entry. Each line is one complete JSON message. It separately describes WebSocket transport as experimental and unsupported, with one JSON-RPC message in each WebSocket text frame. That framing distinction is operationally important: a stdio parser accumulates bytes until a line boundary, while a WebSocket client consumes complete text frames. It is not evidence that either option is faster, more reliable or certified for a production workload.

Local inter-process communication is a boundary design choice rather than permission to improvise the protocol. For example, an organisation might place a same-host broker or operating-system socket between a managed client and a child process. Such an adapter would need to preserve message boundaries, lifecycle ownership and access controls. It would also add another component whose failure and identity rules must be tested. Unless the deployed App Server version explicitly supports the selected local mechanism, do not assume that a native socket form exists merely because the operating system offers it.

A suggested path map contains one row for every boundary crossed:

Hop Owner Reachable by Boundary control Failure to test
Harbour Desk to spawned App Server Harbour Desk process Parent and child processes Inherited standard streams and process lifecycle Child exits or emits a malformed line
Harbour Desk to loopback listener Local service manager Other processes on the workstation Loopback-only bind plus local process controls Unexpected non-loopback bind
Remote client to WebSocket service Deployment operator Hosts able to route to the service TLS and authenticated handshake Unauthenticated or unencrypted connection succeeds

The table is an illustrative threat model, not measured product behaviour. It exposes an important trade-off. A loopback listener avoids external network reachability, but it is broader than a parent-child pipe because other local processes may be able to attempt a connection. A remote listener supports genuinely distributed clients, but adds routing, certificate, authentication, proxy and secret-lifecycle concerns. A local inter-process communication adapter can narrow access through operating-system controls, but increases implementation and testing complexity.

Harbour Desk compares two viable local shapes. In the first, it spawns the App Server and exchanges JSONL over standard streams. In the second, it connects to a listener bound only to the loopback interface. The first provides direct process ownership: Harbour Desk can associate one child with one session and close its streams during shutdown. The second permits process separation and independent restarts, but broadens the boundary from one process relationship to locally reachable processes. Because Harbour Desk has no independent-service requirement, the extra reach does not buy a required capability.

The framing implementation should follow the chosen transport rather than share a permissive “accept anything” parser. For stdio, a suggested client buffers partial reads, extracts complete newline-terminated records, applies a bounded message-size policy chosen by the organisation and parses each record as one JSON value. For WebSockets, a suggested client accepts the documented text-frame model and rejects unexpected binary frames according to its own fail-closed policy. Any local adapter should convert boundaries explicitly; it should not concatenate several messages into ambiguous text or split one JSON value without a defined reassembly rule.

Keep untrusted material out of prompts while testing this plumbing. Use fictional fixtures containing no proprietary source, personal data, credentials or copied terminal output. For example, a test message can carry a fictional repository name and synthetic file path. The useful test is whether one framed message becomes one parsed protocol object, not whether sensitive real-world content passes through the channel.

The decision rule for framing is therefore transport-specific: use line boundaries for the documented stdio JSONL stream, and use one text frame per JSON-RPC message for the documented experimental WebSocket model. Do not add auto-detection that silently treats malformed lines, arbitrary frames or console output as protocol messages. Flexibility at this layer can turn corruption into plausible but incorrect state.

When a local child process is the safer default

A child process and a loopback service are both local, but they are not equivalent trust boundaries. With a child, the client generally owns creation, stream attachment and termination. With a loopback listener, the server has an address and can outlive or accept connections independently of a particular client. Loopback limits routing to the host; it does not mean that only the intended application can connect.

For same-workstation development, keep any listener loopback-only. A wildcard bind, a host’s externally reachable address or a public interface is not an acceptable substitute for loopback testing. If a listener is unnecessary, do not create one. This rule reduces exposure without making unsupported claims that local processes are inherently trustworthy.

A suggested child-process procedure for Harbour Desk is:

  1. Resolve the intended App Server executable through the organisation’s deployment mechanism, rather than accepting an executable path from untrusted input.
  2. Spawn one child with separate protocol input, protocol output and diagnostic handling. Do not merge diagnostic text into the JSONL stream.
  3. Record an internal association between the parent session and the child process identifier. Do not put prompts, source text or credentials into process arguments.
  4. Read complete JSONL records with bounded buffering. Treat parse failures as transport or protocol faults requiring review, rather than guessing where a message ends.
  5. On client shutdown, stop sending work, close the protocol channel and apply an explicitly documented bounded termination policy. A forced termination fallback is an implementation choice, not an App Server delivery guarantee.
  6. Have a human review logs and failure records before using the tool for consequential code, command or approval decisions.

In a fictional test fixture, Harbour Desk starts a child and receives two complete JSONL records split across several operating-system reads. The suggested parser buffers the fragments until each newline appears, then parses the records in order. The expected sample outcome is two protocol objects and no interpretation of an incomplete trailing fragment. This is an example acceptance criterion, not a reported test result.

If Harbour Desk instead needs a separately managed local service, the suggested alternative is a loopback-only listener. Select an available local port through the deployment harness, avoid embedding it as an assumed product endpoint, and verify the actual bound address using operating-system tooling. The client should fail closed if configuration resolves to a non-loopback interface. Port secrecy is not an access-control mechanism.

A compact threat walk-through shows why these details matter:

  • Careless bind address: a developer intends local testing but configures a wildcard listener. The failure occurs at the exposure boundary: another host may now have a route. The corrective decision is to stop the service, restore a loopback-only bind and verify reachability again before protocol testing.
  • Copied token: a credential is pasted into a command, URL, screenshot or diagnostic message. The failure is secret handling, even if the server later rejects the connection. Remove the exposed material, follow the organisation’s credential-rotation procedure and redesign injection so the value never enters those surfaces.
  • Over-broad client: a review helper can start arbitrary servers, connect to arbitrary addresses or silently broaden command scope. The failure is excessive client authority. Constrain destinations and process invocation in configuration, present consequential actions for human review, and reject unrecognised scope rather than inferring consent.

This threat walk-through does not prove that stdio is secure in every environment. A compromised parent, child or workstation can still violate the boundary. Its purpose is to distinguish avoidable network exposure from host-level trust. The practical trade-off is operational independence versus reach: a separately managed listener can survive a client restart, whereas a child is easier to bind to one client lifecycle. Choose the listener only when independent lifecycle management is a documented requirement.

Decide whether remote reachability is actually required

“Remote” means that a legitimate connection must cross a host boundary; it does not mean that attaching from another process is merely convenient. Before accepting that requirement, name the remote operator, source environment, network route, service owner and revocation process. If any of these remains undefined, retain a local transport and resolve the operational design first.

The distinction matters because the documented WebSocket option is experimental and unsupported. As of 6 October 2026, OpenAI’s App Server documentation says it uses one JSON-RPC message per WebSocket text frame. The experimental WebSocket transport can support deliberate evaluation, but it is not a statement of production support, availability, performance or recovery guarantees.

For any non-local connection, require TLS and an authenticated WebSocket handshake. OpenAI documents that the client presents a bearer credential in the Authorization header during the handshake and that authentication is enforced before JSON-RPC initialisation. The same documentation warns that non-loopback WebSocket listeners currently allow unauthenticated connections by default during rollout. Consequently, mere server startup is not evidence of safe remote configuration.

No credential belongs in source code, a URL, a command-line argument, a log, a screenshot, an example repository or a prompt. OpenAI recommends a token file rather than passing a raw bearer token on the command line. In an actual deployment, use an organisation-approved secret source with appropriately constrained file or process access, and redact authentication headers from telemetry. Examples should use fictional stand-ins only and must never resemble a usable secret.

A suggested remote-readiness review is:

  1. Write the business or operational reason that a same-host client cannot perform the task.
  2. Confirm that the deployment path terminates TLS at a controlled component and preserves the authenticated WebSocket handshake according to the deployed design.
  3. Verify that an unauthenticated handshake fails before any JSON-RPC message can be processed.
  4. Verify that a credential reaches the handshake from an approved secret source without appearing in arguments, URLs, logs, screenshots, prompts or repositories.
  5. Check proxy, certificate and operating-system constraints in the actual target environment. Do not infer them from a developer workstation.
  6. Require human review of the exposure map and failure evidence before enabling non-local routing.

Harbour Desk fails the first step: it has no legitimate user or component on another host. It therefore rejects a public listener rather than compensating for unnecessary exposure with more configuration. It may retain a loopback WebSocket fixture solely for local framing tests, but that test arrangement remains loopback-only and does not establish remote suitability.

Before implementing either local option, verify it against the deployed binary rather than a remembered schema. A suggested ordered procedure is:

  1. Generate schemas from the deployed App Server binary. Use the binary’s documented schema-generation facility in the controlled development environment. OpenAI states that each output is specific to the Codex version run, so record the binary version and keep generated artefacts aligned with it.
  2. Start the selected local transport. For Harbour Desk, spawn the stdio child. If evaluating WebSockets locally, bind only to the loopback interface and use synthetic, non-sensitive fixtures.
  3. Prove that no non-local route is present. Inspect the host’s listening sockets and attempt the organisation’s approved route check from a separate test context. For stdio, confirm that no network listener was created. For loopback, confirm that only loopback addresses are bound. A missing browser response is not sufficient evidence.
  4. Document the selected boundary. Record owner, transport, framing, permitted peers, bind scope, credential treatment, startup and shutdown ownership, schema version and reviewer. Keep secrets and untrusted content out of that record.

The acceptance rule is negative as well as positive: the intended local client must be able to communicate, and an unintended non-local peer must have no route. If the second condition cannot be demonstrated, stop before adding event handling or approvals. For a truly remote design, replace “no route” with evidence that only the intended TLS-protected, authenticated route exists; remote enablement still requires human review.

Harbour Desk’s illustrative decision record reads: Harbour Desk will spawn one App Server child and exchange documented JSONL messages through standard input and output because the client owns the process, all required work occurs on one workstation, no independent service lifecycle is needed and a network listener would enlarge the reachable boundary without adding a required capability. A loopback WebSocket may be used only as an isolated local framing fixture; a public listener is rejected. Generated schemas will be tied to the deployed binary version, diagnostics will remain separate from protocol output, and consequential use will remain subject to human review. With that boundary fixed, the next implementation task is the per-connection lifecycle: opening the channel, initialising it exactly once and moving into a ready state.

Turn JSON-RPC Traffic Into Recoverable Client State

A recoverable Codex App Server client should treat the connection, conversation and streamed output as separate state domains. The practical distinction is between transport state, which disappears when a connection closes, and conversation state, which may be addressable again through a retained thread identifier. A third domain, presentation state, can contain incomplete deltas that are useful on screen but are not yet an authoritative record.

OpenAI’s App Server documentation, as accessed on 6 October 2026, defines the required order for each transport connection: send one initialize request before invoking any other method, send the initialized notification, and only then use methods such as thread/start, thread/resume and turn/start. Build that order as a state machine rather than as a convention scattered across event handlers. If the current connection has not completed its initialisation exchange, no conversation method is eligible to leave the client.

This section uses Harbour Desk, a fictional internal review tool, as an illustrative worked example. Its request, identifiers, ledger and recovery messages are examples rather than product guarantees. Harbour Desk sends the prompt “summarise this diff”, records the returned thread identifier, renders item deltas, loses its connection, and then assesses whether to resume. It persists an item only when it receives item/completed. Do not put an actual confidential diff, credentials or other untrusted data into test prompts; use a synthetic fixture and inspect consequential output before acting on it.

Conceptual glowing thread spool passing through circular turns to translucent streaming cards and a solid completed card.
An original conceptual view of provisional event flow becoming authoritative completion state.

Gate every fresh connection through initialisation

Model the connection with at least four client-owned states: disconnected, initializing, ready and closing. These are illustrative names, not documented interface labels. On opening a fresh transport, move to initializing, allocate one request identifier for initialize, and block all thread and turn operations behind that exchange. After a successful response, send the initialized notification and move to ready. A reconnect creates a new transport connection and therefore requires a new, single initialize request; the fact that an earlier connection was ready does not initialise its replacement.

An illustrative dispatcher can enforce the boundary without allowing callers to bypass it:

onConnectionOpen(connection):
    assert connection.state == "disconnected"
    connection.state = "initializing"
    sendRequest(connection, "initialize", localCapabilities)

onInitializeResponse(connection, response):
    validateAgainstDeployedSchema(response)
    sendNotification(connection, "initialized")
    connection.state = "ready"
    releaseQueuedConversationIntent()

sendConversationRequest(method, parameters):
    if activeConnection.state != "ready":
        queueIntentOrReturnLocalError(method)
        return
    sendRequest(activeConnection, method, parameters)

This pseudocode is a suggested client structure, not a statement about an official library. Validate message shapes against schemas generated from the deployed codex app-server binary because the official documentation says those outputs are specific to the Codex version being run. The practical procedure is to generate the schemas in the controlled development environment, bind the client’s decoder to that version, and fail closed on a response that cannot be associated with the outstanding request. Human review is required before a schema change alters behaviour that can run commands, modify files or affect another consequential workflow.

Do not interpret “send a single initialize” as permission to send it repeatedly until one succeeds. Keep an explicit per-connection flag such as initializeSent. If a second component attempts initialisation on the same live connection, stop it locally and report a client defect. If the connection ends before initialisation completes, discard that connection’s pending exchange and initialise once on the newly opened connection.

After the initialized notification, choose one conversation operation. For a new conversation, issue thread/start and retain the returned thread identifier before starting its first turn. For an existing conversation selected for recovery, issue thread/resume with the retained thread identifier. Only after the thread is started or resumed should the client issue turn/start. This makes the complete sequence auditable:

  1. Open a transport connection.
  2. Send exactly one initialize request on that connection.
  3. Validate its response and send initialized.
  4. Start a new thread or resume a retained thread.
  5. Start a turn within that thread.
  6. Reduce incoming item events into provisional and authoritative state.

Add a negative protocol test rather than relying only on the happy path. On an isolated local test instance, deliberately send a harmless request before initialize. The expected observation, according to the App Server documentation accessed on 6 October 2026, is a Not initialized error. Record whether the client surfaces that error and refuses to infer that the conversation operation succeeded. This is a test of documented behaviour, not a broader guarantee about every malformed sequence, connection lifetime or future version. Restore the ordinary sequence immediately afterwards by opening a fresh test connection.

The pass rule for this test is not merely “an error appeared”. Pass only when the premature request creates no local thread or turn record, the error remains associated with its request identifier, and the following valid test starts from a newly initialised connection. A timeout, dropped connection or different response is not equivalent to the documented error; preserve the evidence and investigate version and schema alignment rather than rewriting local state to fit an assumption.

Reduce threads, turns and items without confusing their lifetimes

The documented object hierarchy provides the reducer boundary: a thread is a conversation and contains turns; a turn represents one user request plus the agent work that follows; turns contain items and incremental updates. Consequently, a thread identifier must not double as a turn identifier, and neither should be used as an item key. Use a composite lookup such as (threadId, turnId, itemId) for item state, while retaining each identifier in its own typed field.

In the fictional Harbour Desk example, the operator chooses a synthetic diff fixture and asks, “summarise this diff”. After the connection is initialised, Harbour Desk calls thread/start. It receives a thread identifier, validates its shape, and stores it before issuing turn/start. The returned turn identifier is attached to that request only. When an item begins and later emits deltas, Harbour Desk associates those events with both the active turn and the item identifier rather than appending text to a thread-wide buffer.

The distinction matters when more than one turn exists. A thread can survive beyond a single request, but a turn should represent only one request and the resulting work. If Harbour Desk later asks “list the risky changes”, that is a fresh turn in the same thread, not a continuation of the earlier turn identifier. The client may show both turns in one conversation view, but its reducer must not merge their item streams. The decision rule is: reuse a retained thread only when conversational continuity is intended; always bind each new user request to the turn created for that request.

Maintain a labelled state ledger. The following prose ledger is an illustrative implementation record, not an App Server response format:

  • Connection state: ready while the current connection has completed its own initialisation sequence; change it immediately when closure is observed.
  • Thread identifier (ID)A value used to distinguish one record, task, source or object from another. Open glossary entry: the validated identifier returned for Harbour Desk’s conversation; retain it separately so that a later recovery assessment can consider thread/resume.
  • Turn ID: the identifier for “summarise this diff”; use it to scope events and any subsequent operator decision to this request.
  • Item ID: the identifier of the item currently producing updates; never infer it from arrival order.
  • Provisional display: the latest rendered composition of received deltas, explicitly marked incomplete and held outside the durable completed-item store.
  • Authoritative persisted item: empty until a valid item/completed event supplies the final item; replace or populate it from that event rather than from the display buffer.

Apply every event only after checking that its identifiers refer to a known thread and turn. An event for an unknown identifier should enter a bounded diagnostic path, not create a plausible-looking conversation automatically. Similarly, when an event is structurally invalid against the deployed schema, quarantine its metadata without recording sensitive payloads and stop the affected reduction. Do not place secrets, private source text or untrusted event content in diagnostic prompts, logs, screenshots or example repositories.

The ledger also clarifies deletion and retention. Closing a connection removes connection-specific request correlation and readiness state, but it need not erase the retained thread identifier or a previously completed item. Ending a turn stops it from being the active target, but does not make its items belong to the next turn. By contrast, provisional display text may be discarded whenever its stream becomes ambiguous. Prefer losing an incomplete visual fragment to promoting it into a durable business record.

The Codex SDK is adjacent to this design, not its implementation target. OpenAI’s SDK documentation, accessed on 6 October 2026, directs automated coding tasks, including CI jobs, to the SDK, while directing custom clients that handle conversation history, approvals and streamed agent events to App Server. Harbour Desk is managing the latter stateful protocol surface. A CI worker that submits bounded jobs should follow the SDK route rather than adopting this client reducer merely because both can initiate coding work.

Render progress without mistaking deltas for a record

A delta is useful for responsiveness, but it is not the completed item. OpenAI’s App Server documentation says that item/completed sends the final item and should be treated as authoritative. It also warns that a final plan item may not exactly equal concatenated deltas. Therefore, keep two representations: an ephemeral view model assembled for live rendering, and a durable item record populated from the completion event.

For Harbour Desk, suppose the active item emits several text fragments before the connection fails. The client may render those fragments beneath an “in progress” treatment, but it must not save their concatenation as the final summary, export it, or use it as an input to a consequential decision. When item/completed arrives, validate its identifiers and final item, then replace the item’s provisional display with the authoritative content and persist that completed object. This replacement procedure is essential even if the final text appears similar.

An illustrative reducer separates the paths:

onItemDelta(event):
    key = requireKnownKey(event.threadId, event.turnId, event.itemId)
    provisional[key] = applyDisplayDelta(provisional[key], event)
    renderAsIncomplete(provisional[key])

onItemCompleted(event):
    key = requireKnownKey(event.threadId, event.turnId, event.item.id)
    finalItem = validateCompletedItem(event.item)
    persistedCompletedItems[key] = finalItem
    provisional.remove(key)
    renderAsCompleted(finalItem)

The sample deliberately does not derive persistedCompletedItems from provisional. It also does not promise that a given sequence of deltas is complete, replayable or unique. The client should correlate by documented identifiers and validate the completed event, while treating arrival order alone as insufficient evidence.

Define persistence narrowly. It is reasonable to retain connection-independent identifiers, validated completed items and the client’s recovery status. It is not reasonable to label an interrupted delta buffer “complete” merely to simplify storage. If a product must retain the provisional buffer for troubleshooting, store it in a separately labelled, access-controlled transient area with a clear incomplete status and an appropriate deletion policy; do not mix it with completed conversation records.

Test this distinction with a fictional fixture containing predictable, non-sensitive text. Feed several deltas into the reducer and verify that the live view changes while the authoritative item field remains empty. Then feed a valid item/completed whose final content differs from the concatenated fragments. The expected client outcome is that the persisted record equals the completed item and the provisional buffer is removed. This expected outcome tests the suggested reducer design; it is not a benchmark or a promise about what content the model will produce.

The decision rule for downstream consumers is equally strict: a component that needs progress may subscribe to provisional state and display its incomplete status; a component that exports, indexes, audits or acts on an item must consume only the validated completion record. Any action with financial, operational, access-control or code-execution consequences requires human review even after completion, because authoritative protocol state does not establish that the content is correct or safe.

Resume deliberately after a broken stream

A broken connection creates uncertainty, not proof that the turn failed or succeeded. The official sources do not promise exactly-once replay, automatic duplicate suppression or guaranteed recovery after disconnection. Reconnection policy and deduplication are client choices. Preserve authoritative completion state already received, mark unresolved items as interrupted, and assess whether a retained thread is suitable for resumption rather than silently repeating the user’s request.

Harbour Desk loses connectivity after rendering provisional deltas for “summarise this diff”, but before receiving item/completed. Its ledger now reads: connection state disconnected; retained thread ID present; turn ID present but unresolved; item ID present; provisional display marked interrupted; authoritative persisted item empty. It does not convert the visible fragments into a final summary and does not immediately send another turn/start.

Use explicit resume criteria. A suggested client may attempt thread/resume only when all of the following are true: the reconnect has completed a fresh initialize/initialized exchange; the retained thread identifier was obtained from a validated response; the operator still intends to continue that conversation; completed items already held locally remain immutable; and the client is prepared to reconcile incoming events without assuming replay uniqueness. If any criterion fails, stop and request human review or begin a clearly separate conversation.

The procedure is: close and discard the failed connection’s request table; retain validated completed items and identifiers; open a new connection; send exactly one new initialize; send initialized after its successful response; issue thread/resume with the retained thread identifier; and inspect subsequent state before deciding whether another turn is needed. Do not resend “summarise this diff” merely because the original display was incomplete. A repeated request could represent additional work rather than recovery.

Show an operator-facing message that states uncertainty without claiming server behaviour. For example: The live stream was interrupted before this item was confirmed complete. Harbour Desk has retained the conversation identifier and any earlier completed items. Reconnect and assess the resumed thread before starting another turn. The incomplete preview will not be saved as the final item. This is sample wording, not an official interface message.

If a completed event is subsequently received and validates against the known thread, turn and item, use it as the authoritative persisted item and discard the interrupted preview. If no completion can be established, retain the status as unresolved. Where the client observes an event corresponding to an item it has already persisted, compare stable identifiers and validated content under a documented local policy; do not claim the server suppressed a duplicate. Material conflicts must be presented to a human rather than resolved by arrival order.

Run a connectivity interruption test during deltas. With a synthetic prompt and fixture, initialise the connection, start a thread and turn, wait until provisional item updates are being rendered, then interrupt connectivity in the controlled test harness. The expected observation is not guaranteed continuation. It is a deliberate resume assessment: the client marks the connection unavailable, preserves prior authoritative completions, leaves the interrupted item unresolved, performs fresh initialisation after reconnect, and evaluates thread/resume against the criteria above.

Pass the interruption test only if provisional text remains separate from persisted completion state, no automatic duplicate turn is started, and the operator can see what is known and unknown. Also confirm that the retained thread identifier survives while old connection readiness and pending request correlations do not. A later item/completed may resolve the item; absence of that event must not be rewritten as completion.

Finally, prepare the reducer for the next protocol gate. The documentation describes approvals as server-initiated JSON-RPC requests containing thread and turn identifiers. Before a human chooses among scoped possibilities such as accept, acceptForSession, decline or cancel, the client must present the command, scope and matching conversation context. Do not auto-approve. The state work above ensures an approval cannot be attached merely to whichever conversation happens to be visible.

  • Open each fresh connection in a non-ready state; send exactly one initialize, then initialized.
  • Start or resume a thread only after initialisation; start a turn only after establishing the thread.
  • Keep connection, thread, turn and item identifiers in separate typed fields.
  • Render deltas as provisional progress; persist only a validated item/completed item.
  • After interruption, preserve completed state, mark incomplete displays unresolved and assess resumption deliberately.
  • Do not assume exactly-once replay or server-side duplicate suppression.
  • Run both the pre-initialisation negative test and the interruption-during-deltas test with synthetic data.
  • Before the approval gate, verify that command, scope, thread and turn context can be shown to a human reviewer.

Make the Approval Request a Real Human Control

An approval request is not an informational event and should not pass through the same rendering path as streamed progress. OpenAI’s Codex App Server documentation, fetched on 6 October 2026, describes the exchange directly: “The app-server sends a server-initiated JSON-RPC request to the client, and the client responds with a decision payload.” The practical distinction is that the server is awaiting a scoped decision; the client is not merely acknowledging receipt. A client that silently accepts, derives consent from a timeout or treats the request as another status update removes the human gate that the protocol enables.

Implement approvals as a separate state machine with at least four client-side states: received, awaiting human review, decided and closed. These are illustrative application states, not documented protocol labels. On receipt, validate the request structure against schemas generated from the deployed codex app-server version, extract its supplied thread and turn identifiers, and associate it with exactly one visible review record. Do not insert untrusted command text, tool output or prompts into another model request to decide whether approval is safe. Keep secrets and other untrusted data out of prompts.

Respond only after an identified person has seen the proposed command and scope in the context of the matching active thread and turn. If the client cannot establish that match, the request must not be actionable. A conservative client may leave the request unresolved while it obtains fresh state, or permit the reviewer to decline or cancel it, but it must not manufacture consent. This sacrifices some convenience when state is incomplete in exchange for preventing an approval from crossing conversation boundaries.

For example, an illustrative internal client called Harbour Desk could receive an approval while its fictional review helper is processing a turn. Harbour Desk would pause that approval record in an awaiting-review state without claiming that the whole turn has stopped. It would then show the named reviewer a redacted command summary, the thread identifier, the turn identifier and the proposed approval scope. Only that person’s deliberate choice would produce the decision payload. Harbour Desk, its names and its behaviour in this guide are fictional implementation examples, not product guarantees.

Teal light paths pass through a transparent glass vessel and a brass latch; an outstretched hand pauses above clear beads, suggesting bounded recovery and human review.
An original conceptual view of a bounded recovery path with a separate human-control boundary.

Display scope before a person chooses

Thread and turn identifiers answer a different question from the command summary. The summary explains what the requested operation appears to do; threadId and turnId identify the conversation and unit of work to which the request belongs. The App Server documentation says that approval requests include those identifiers and instructs clients to use them to scope interface state to the active conversation. A familiar-looking command is therefore insufficient evidence that the request belongs in the panel currently on screen.

Build the panel from validated request data rather than from whichever conversation happens to be visible when rendering completes. An illustrative Harbour Desk panel might present the following, with values deliberately fictional and redacted:

Reviewer
Priya N. — authenticated human reviewer
Command summary
Proposes running a package inspection command; arguments and paths redacted
Thread
thread-harbour-17
Turn
turn-review-04
Proposed scope
This request only, or the documented session scope if the reviewer deliberately selects session acceptance
Available decisions
Accept, session acceptance, decline or cancel

Those labels represent the documented decision possibilities accept, acceptForSession, decline and cancel; they do not create permissions beyond their documented scope. In particular, do not describe session acceptance as permanent access, workspace-wide authority or permission for unrelated threads. The client should explain the proposed scope without extrapolating broader rights. It should also avoid preselecting any decision, making one button visually inevitable, or treating keyboard focus as consent.

Use this rendering procedure. First, capture the active panel’s thread and turn from authoritative client state. Secondly, compare both identifiers with the approval request. Thirdly, produce a redacted description that conveys the operation’s purpose without retaining code, file paths, prompts, credentials or complete output. Fourthly, display the scope beside—not behind—the choices. Fifthly, require an intentional action from the named reviewer and bind the resulting decision to the pending request that was displayed. Finally, close that review record so a delayed click cannot decide a replacement request.

The negative test is mandatory: fabricate an illustrative approval fixture carrying thread-other-92 while Harbour Desk displays thread-harbour-17. The foreign request must not be actionable in the active panel. It may appear in a quarantined diagnostic view containing only redacted metadata, but none of its accept, session-acceptance, decline or cancel controls should be bound to the active conversation. Repeat the test with a matching thread but a different turn. The decision rule is exact equality for both identifiers; similarity, recent activity and matching command text are not substitutes.

This strict binding adds interface work when several turns are visible, but it makes the authority boundary inspectable. A multi-thread client may maintain separate approval queues per thread; a simpler client may permit review only for the foreground thread. Either design is defensible if a request from another thread cannot inherit the foreground panel’s reviewer action. For consequential operations, retain human review even when the command summary appears routine.

Keep approval timing separate from sandbox capability

An approval asks when a human permits a proposed action to proceed; the sandbox determines what the agent can do if execution is attempted. These controls are related but not interchangeable. OpenAI’s agent approvals and security documentation, fetched on 6 October 2026, states that network access is turned off by default. An approval decision does not itself prove that network access, filesystem reach or another capability is available, and a sandbox capability does not mean a person has approved its use at this point in the turn.

Model these as two independent checks in the client’s explanation. One line should describe the proposed approval scope; another may describe known execution constraints, provided those constraints come from current, trustworthy configuration rather than an inference from the decision label. Do not render “accept” as “grant network access”, and do not render a network-disabled baseline as “approval unnecessary”. The app server’s decision vocabulary should be passed with its documented meaning, not repurposed as a general permission system.

A practical review sequence is: identify the requested operation; show its redacted command summary and supplied scope; show the matching thread and turn; state any independently verified sandbox constraint; then ask the human to decide. After a decision, let the protocol and execution environment determine what can happen. The panel should not predict success merely because the reviewer accepted, nor claim execution merely because it sent a response.

For example, suppose Harbour Desk shows a command that appears to contact an external package service while the applicable sandbox has network access off. The reviewer still evaluates the request’s purpose and scope. Declining answers the consent question; network-off answers the capability question. If the reviewer accepts instead, the client must not announce that external access has been granted or completed. The meaningful trade-off is between concise presentation and preserving this distinction: adding a separate capability line takes space, but avoids turning one control into a misleading proxy for the other.

Verify the separation with an illustrative two-by-two test: approval pending with network off; approval declined with network off; approval accepted with network off; and approval pending under a separately configured capability set. Review the panel wording in each case. The approval choices must not change merely because capability differs, and the capability description must not report consent before a person decides. Use test fixtures rather than real secrets, commands or sensitive prompts. A human must review any consequential conclusion about whether the resulting deployment policy is appropriate.

Exercise decline and cancel as safety tests

Decline and cancel are separate documented decision possibilities, but neither should be presented as execution. Avoid inventing a universal semantic difference beyond the deployed schema and documentation: the client should preserve the exact selected value, close the matching review state and observe subsequent server events without relabelling either outcome as success. This is particularly important where a local interface uses friendly wording; the stored protocol decision must remain unambiguous.

Work through a controlled example. Harbour Desk’s fictional review helper is processing turn-review-04 in thread-harbour-17 when an approval request appears. Priya compares the displayed thread and turn with the active conversation, examines the redacted command summary and proposed scope, and selects decline. The client binds that choice to the pending request, sends the corresponding decision payload, marks the review record closed and does not create any local “executed” state. The operator then repeats the fixture with a fresh request and selects cancel. Harbour Desk again closes only the matching review record and does not interpret the choice as command execution.

This is a verification exercise, not evidence that a particular real command was blocked. Confirm outcomes from protocol state and authoritative subsequent events rather than from the absence of visible output. A quiet terminal, unchanged file or short wait is only an observation and is not a reliable execution test. The client should never fabricate a completed item, successful command or failed command solely from an approval response.

Use the following concrete test procedure:

  1. Create a fictional fixture with a valid-looking request shape, redacted command summary, active threadId and active turnId. Validate it against schemas generated for the deployed App Server version.
  2. Open it in the approval panel and confirm that no choice is preselected and no timer produces acceptance.
  3. Select decline as the named test reviewer. Confirm that exactly the pending review record is closed and that the client does not mark the command or turn as executed.
  4. Submit a new fixture for the same active turn, then select cancel. Confirm the same non-execution rule while preserving cancel as the selected decision rather than rewriting it as decline.
  5. Create a request for the active thread but an earlier, already-closed turn. Confirm that the client rejects it as stale for interactive approval, exposes no acceptance action and does not attach it to the current turn.
  6. Create a request from another thread. Confirm it cannot be acted upon in the active panel, even if its redacted command summary is identical.
  7. Inspect emitted telemetry and verify that it contains no command text, code, file paths, prompts, credentials or full tool output.

Stale-turn rejection is a client safety policy, not a claimed server delivery guarantee. Define “stale” using the client’s authoritative turn lifecycle: for example, a turn recorded as closed must not regain an actionable approval control merely because a delayed request is rendered. If the client cannot establish the turn’s status after an interruption, put the request into a non-actionable reconciliation state and require fresh human review after state is restored. The trade-off is that a legitimate delayed request may need to be requested again, but this is preferable to attaching old consent to new work.

Also test repeated clicks and delayed interface callbacks. Once one decision has been committed for a pending request, disable all associated controls and ignore later local callbacks for that review record. This is client-side defensive behaviour, not a promise of duplicate suppression by the service. Do not automatically transform a timeout, closed browser tab or lost connection into accept, session acceptance, decline or cancel; an absent decision is not consent.

Record a privacy-minimised decision trail

An operational decision trail should prove that scoping and human review occurred without becoming a second repository for sensitive work. Distinguish decision metadata from request content. The useful metadata is the minimum needed to reconstruct which client state was reviewed: a pseudonymous reviewer reference, a locally assigned review-record identifier, redacted or one-way-correlated thread and turn references, the decision selected, the proposed scope category, timestamps, and a reason code such as scope mismatch or stale turn. The command, source code, file paths, prompt, credentials and full tool output should remain absent.

An illustrative audit entry could read: “Review record rv-104; reviewer person-27; thread reference th…17; turn reference tu…04; scope category ‘single request’; decision ‘decline’; reason ‘scope not expected’; content retained: none.” A cancellation entry might use the same limited fields with decision “cancel”. This sample is a fictional format, not an App Server response or a prescribed compliance record. Whether identifiers should be truncated, hashed or stored in a restricted correlation table depends on the organisation’s operational need and privacy review.

Apply data minimisation before the telemetry sink, not merely in a dashboard. Construct audit events from an allowlist of fields rather than serialising the approval request and attempting to remove sensitive properties afterwards. Reject unexpected fields at the logging boundary. Keep diagnostic exceptions from dumping the original JSON-RPC object, and ensure screenshots used in defect reports contain no credentials, paths, prompts or command details. Never put an approval payload into a model prompt for summarisation.

Verify redaction with synthetic canaries. Put unmistakably fictional strings into the fixture’s command, path, prompt and output fields, then exercise accept, session acceptance, decline, cancel, scope mismatch and stale-turn handling. Search the permitted telemetry export for every canary and inspect structured fields as well as exception messages. The expected rule is absence: if any canary survives, block the telemetry change until a human reviewer has checked the correction. This procedure demonstrates the redaction path with test data; it does not guarantee that every future field will be harmless.

Retention presents a genuine trade-off. More detailed records can simplify debugging, but they increase the consequences of accidental disclosure and can obscure the specific fact being audited: who reviewed which scoped request and what decision they made. Prefer a narrow, documented retention purpose and access boundary. Consequential choices about identity, retention and employee monitoring require human privacy, security and legal review rather than being inferred from this implementation example.

Finally, preserve four client invariants across any reconnect: the active-thread scope must still bind approvals to the correct thread and turn; each pending or closed review state must remain distinguishable; redaction must continue before display exports and telemetry; and no request may become actionable without an explicit human decision. Reconnection and deduplication are client policies, not delivery guarantees. If restored state cannot prove those invariants, keep the approval non-actionable until a person reviews reconciled state.

Rollout warning: according to the Codex App Server documentation, non-loopback WebSocket listeners may allow unauthenticated connections by default during rollout. Do not interpret successful remote connectivity as evidence that authentication is active. Local tests should bind only to 127.0.0.1. A genuinely remote design requires wss://, TLS and an authenticated WebSocket handshake before any JSON-RPC initialize request is accepted.

Expose Remote Access Only With Tested Failure Containment

The Codex App Server is the documented surface for custom clients that manage authentication, conversation history, approvals and streamed agent events. The Codex SDK is instead directed towards automated coding tasks, including continuous integration jobs. This release gate therefore concerns a stateful custom client, not a continuous integration deployment pattern or a general comparison of integration options.

Transport choice changes both the trust boundary and the maturity of the connection. The app-server documentation describes standard input/output as the default transport, carrying newline-delimited JSON, while WebSocket transport sends one JSON-RPC message per text frame and is experimental and unsupported. As of 6 October 2026, that status provides no production-support guarantee. The release decision is consequently not “does WebSocket work?” but “is remote reachability necessary, authenticated before protocol traffic, and contained when a dependency fails?”

Use this sequence for the fictional Harbour Desk release check:

  1. Keep developer and same-host tests on 127.0.0.1; do not bind a test listener to a non-loopback interface for convenience.
  2. Record why a child process using the default standard input/output transport cannot satisfy the design. Remote WebSocket exposure should require a concrete cross-host need.
  3. For an approved remote design, require wss://, TLS and bearer authentication in the WebSocket handshake.
  4. Prove that a handshake without authentication is rejected before JSON-RPC begins.
  5. Inspect the deployed binary’s version-specific documentation or generated schema before defining readiness and protocol checks.
  6. Exercise overload, proxy denial and interrupted-request paths without assuming replay or duplicate suppression.
  7. Submit the evidence and remaining risks for local organisational and human review before release.

The decision rule is strict: if Harbour Desk cannot prove pre-JSON-RPC authentication, prevent credential disclosure, or bound retries, retain a loopback-only design. TLS protects a remote connection in transit, whereas authentication decides whether the peer may begin a session; one is not a substitute for the other. All remote controls must also pass the organisation’s own architecture, operations and security review.

Require authentication before JSON-RPC begins

The documented connection order has two separate gates. First, the client presents Authorization: Bearer <token> during the WebSocket handshake, and app-server enforces authentication before JSON-RPC initialize. Second, on every accepted transport connection, the client sends one initialize request before any other method and then sends an initialized notification. Requests sent before initialisation receive a Not initialized error. Authentication establishes who may connect; initialisation establishes protocol readiness.

An illustrative Harbour Desk connection state machine is:

DISCONNECTED
  -> TLS_CONNECTING
  -> AUTHENTICATING_WEBSOCKET
  -> INITIALIZING_JSON_RPC
  -> READY
  -> CLOSING or DISCONNECTED

This is sample client design, not a product guarantee. The client should allow no application request in the first four states. A socket-open event alone must not enable conversation controls, because the handshake or JSON-RPC initialisation may still fail.

Run the negative authentication test with a non-secret test identity and an isolated environment:

  1. Attempt the remote WebSocket handshake without an Authorization value.
  2. Verify that the connection is rejected before the client can send initialize.
  3. Repeat with malformed or inapplicable test credentials, without recording their values.
  4. Confirm that failure telemetry contains only a redacted reason, timestamp and test correlation identifier.
  5. Only then use an authorised test credential and verify the required initialize, followed by initialized, on that new connection.

Do not replace this test with a post-connection method failure. Rejection after initialize would place the control at the wrong boundary. Conversely, a Not initialized response is useful for testing protocol order but does not demonstrate authentication. Release only when both rejection paths are distinguishable and tested.

Harbour Desk must also keep approval handling behind the ready state. An approval is a server-initiated JSON-RPC request carrying thread and turn identifiers. The client should display the proposed command and its scope before a human returns a scoped decision such as accept, acceptForSession, decline or cancel. These are decision possibilities, not grounds for automatic acceptance. Consequential actions require human review.

Keep the bearer credential out of the process trail

The app-server documentation prefers --ws-token-file to putting a raw bearer token on the command line. This distinction matters because a protected file can be supplied without embedding the credential in process arguments. A raw command argument, URL or source file can expose the value beyond the handshake boundary. Do not place a credential in code, a URL, a command-line argument, a log, a screenshot, a prompt or an example repository.

For Harbour Desk, use this illustrative handling procedure:

  1. Provision the credential through an organisation-approved secret process outside the application repository.
  2. Write it to a protected token file using the deployment environment’s reviewed access controls; do not print the value during provisioning.
  3. Configure app-server to read that file rather than receiving the raw value as an argument.
  4. Have the client obtain its handshake credential through an equally protected local mechanism, without copying it into diagnostic context.
  5. Redact request headers before any structured logging, exception capture or tracing.
  6. Remove test credentials under the organisation’s established lifecycle procedure when the release check ends.

This procedure deliberately does not prescribe a certificate authority, token-issuance application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry, file mode or secret manager: those details are not established by the assigned app-server sources and depend on the deployment environment. The release evidence should identify the chosen local controls without presenting them as Codex guarantees.

The Harbour Desk secret-leak check should scan application logs, proxy logs, crash output and process listings for the test bearer value. Inspect configuration bundles and the proposed repository contents as well. The expected release condition is absence of the value, but the test report must not reproduce it as evidence. Record only which surfaces were examined and whether remediation remains open.

The trade-off is operational complexity versus disclosure risk. A protected token file requires controlled provisioning and clean-up, but a raw process argument creates an avoidable process trail. The decision rule is therefore categorical: if the deployed command line, URL, repository or routine telemetry contains the bearer value, stop the release and replace that path rather than adding a masking promise afterwards.

Probe readiness and prove rejection paths

A transport connection, an authenticated handshake and an initialised JSON-RPC session are different states. A readiness check must correspond to the deployed binary’s documented behaviour; it must not infer readiness merely because a port accepts connections. The official documentation also notes that generated outputs are specific to the Codex version run. Generate and inspect the relevant schemas or documentation from the exact deployed codex app-server binary before implementing the probe and client validation.

Harbour Desk’s illustrative readiness procedure is:

  1. Pin the binary selected for the release candidate and generate its version-specific protocol artefacts.
  2. Identify only the readiness behaviour documented for that binary. Do not invent a path, status code or response body.
  3. Configure the probe according to that documented behaviour and ensure its output contains no credential, prompt, command or agent response.
  4. Test startup, normal readiness, shutdown and an unavailable dependency relevant to the deployment.
  5. Verify that the deployment does not route ordinary work merely because the network socket is reachable.

A sample acceptance record might say, “Readiness checked against the artefact generated from the deployed binary; unauthenticated WebSocket handshake rejected; application traffic remained disabled until initialisation completed.” This is an example record format, not a claimed test result or specified app-server response.

Test a proxy-denial path separately. With network access denied or the selected destination denied by the organisation’s proxy policy, initiate a reviewed test request that would require that destination. Confirm that the denial remains visible to the human operator and does not cause the client to broaden access, approve a command silently or loop indefinitely. Keep test data non-sensitive and exclude secrets from prompts.

The network sandbox and approval mechanism are separate controls. The security documentation says network access is off by default, while an approval determines whether a proposed action receives human consent at the relevant time and scope. Approval cannot prove that a destination is permitted, and proxy permission cannot stand in for command approval. Harbour Desk should test both controls and preserve that distinction in operator messages.

Destination policy itself is outside this release gate; apply your organisation’s network-governance guidance rather than expanding this check into a proxy configuration tutorial.

The decision rule is to fail closed when the probe contract is unknown, authentication is bypassed, or proxy denial is converted into an unreviewed fallback. The trade-off is that conservative gating may delay availability, but treating reachability as readiness can admit work before the required controls are established.

Use bounded jittered backoff when ingress is full

The documented overload signal is precise: when request ingress is full, app-server rejects new requests with JSON-RPC error code -32001 and the message Server overloaded; retry later. This is narrower than an arbitrary connection error. Apply overload backoff only after matching that documented code and message; route authentication failures, malformed requests, proxy denials and disconnections to their own handling paths.

The documentation recommends an exponentially increasing delay with jitter. The client must add its own cap, retry ceiling, operator visibility and escalation policy. One illustrative Harbour Desk policy is a maximum of five retry decisions, with nominal delays of 500 milliseconds, one second, two seconds, four seconds and eight seconds, each randomised between zero and that attempt’s nominal ceiling. These values are a suggested example, not documented defaults or performance guidance.

if error.code == -32001
   and error.message == "Server overloaded; retry later.":
    if attempts >= 5:
        mark_needs_operator_review()
        stop_retrying()
    else:
        ceiling = min(0.5 * (2 ** attempts), 8.0)
        wait(random_between(0, ceiling))
        retry_only_if_request_is_safe_to_retry()
else:
    handle_as_non_overload_failure()

The final safety check is essential. A retryable overload rejection indicates that ingress rejected the new request, but an interrupted connection may leave the client uncertain whether a different request was accepted or completed. Do not reuse the overload branch as a universal reconnect algorithm.

Test jitter without claiming timing performance:

  1. Use an illustrative test fixture that returns exactly -32001 and Server overloaded; retry later. for a bounded sequence.
  2. Inject a deterministic pseudo-random source so the test can inspect chosen delays without relying on wall-clock anecdotes.
  3. Assert that each delay remains within its attempt ceiling and that later nominal ceilings grow only to the configured cap.
  4. Assert that retrying stops at the configured ceiling and creates an operator-visible state.
  5. Return a different error and verify that overload retry does not run.

Operator visibility should identify the affected request correlation identifier, attempt count and next permitted action, but omit prompts, bearer values and sensitive output. After the retry ceiling, stop automatically resubmitting and ask an operator to review service state and request ambiguity. The trade-off is slower recovery during transient pressure in exchange for avoiding synchronised retry storms and unbounded load.

Ship a recovery runbook rather than a reliability promise

Reconnection and deduplication are client policies, not app-server delivery guarantees. Experimental WebSockets carry no production guarantee, and the sources do not establish exactly-once replay, automatic duplicate suppression or certain recovery after interruption. A release runbook must therefore preserve known authoritative state and expose ambiguity rather than promising that every interrupted request can be replayed safely.

Use thread, turn and item identities separately. A thread contains turns; a turn represents one user request and the following agent work; turns contain items and incremental updates. Retaining a thread identifier supports a deliberate resume design, but it does not prove what happened to an in-flight turn. Render item deltas as provisional progress. When item/completed arrives, persist its final item as authoritative because the documentation warns that a final plan item may differ from concatenated deltas.

An illustrative Harbour Desk interruption runbook is:

  1. Freeze new user submissions and mark the active transport disconnected.
  2. Retain the last authoritative completed items, thread identifier and locally known turn state. Do not promote accumulated deltas to final records.
  3. Open a fresh authenticated transport, then perform the required one-time initialize and initialized sequence for that connection.
  4. Assess whether the documented thread-resume operation is appropriate for the retained thread.
  5. Classify the interrupted request as known complete, known rejected or uncertain from available authoritative events.
  6. Do not replay an uncertain consequential request automatically. Present its command, scope and known state for human review.
  7. Resume rendering from authoritative state and treat subsequent deltas as provisional until completion events arrive.

For example, if Harbour Desk displays two plan deltas and then loses the connection, it may retain those fragments for diagnostic presentation but must not store their concatenation as the final plan. If a later authoritative completed item is obtained through the documented resumed flow, that completed item replaces provisional display state. This illustrates reducer behaviour; it does not guarantee that the server will replay the missing event.

The same caution applies to approvals. After reconnecting, scope every server-initiated approval to its supplied thread and turn identifiers, display the command and scope, and let a human choose among acceptance, session-scoped acceptance, decline or cancel. Never infer approval from a prior transport session, and never put untrusted event data or secrets into a prompt used to make the decision.

The release packet should contain the negative handshake evidence, secret-scan scope, binary-specific readiness basis, bounded overload policy, proxy-denial result and interruption assessment. It should also name the person or team that reviews uncertain consequential requests after automatic retries stop. Release only if the runbook prefers an explicit unresolved state to speculative replay. That choice may require manual intervention, but it avoids converting incomplete event history into a false reliability claim.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this