Codex Security Cloud in Research Preview: A Bounded Pilot for GitHub Scans, Evidence, and Draft PRs

Conceptual illustration of a bounded repository-security pilot

What OpenAI documented, and the decision facing security leads

Your job is to decide whether a small, authorised repository-security pilot deserves engineering time, cloud access and a reviewer’s attention. This is a documentation-led news analysis, reviewed as of 11 October 2026, not a hands-on benchmark. It separates the announcement and current product instructions from our proposed operating rules, so you can approve an experiment without accidentally approving broad access or unattended remediation.

Conceptual illustration of a bounded repository-security pilot
Conceptual illustration of a bounded repository-security pilot. Original conceptual artwork, not a product screenshot or evidence of testing.

OpenAI’s 29 September 2026 DevDay recap announced Codex Security Cloud for scanning GitHub repositories, investigating findings and preparing fixes in the cloud. As reviewed on 11 October 2026, the current product overview still calls the Cloud plugin a research preview on the web and in the desktop app. [DevDay 2026 Recap; Codex Security]

The DevDay recap names Pro, Business, Enterprise and Edu access on desktop and web. The current help page also lists these four plan families. The Cloud overview additionally requires a workspace with Codex Security Cloud access, a connected GitHub repository and a compatible Codex Cloud environment; it directs users without access to their workspace administrator. [DevDay 2026 Recap; Codex Security; Codex Security]

The decision should be narrower than “adopt a security agent”. A useful first question is whether the evidence produced for one familiar repository can support a better human decision about a suspected defect. Define that decision before choosing a scan configuration. Otherwise, the team may collect interesting reports without learning whether anyone can investigate, reject or safely act on them within the available review time.

Throughout this article, artificial intelligence (AI)Computer systems designed to perform tasks that normally require human intelligence, such as understanding language, recognising patterns, or making predictions. Open glossary entry output means material awaiting evaluation, not an independent authority to change a service. A pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry is the review object discussed later. We use the singular abbreviation only where it helps distinguish that object from a scan or finding. The proposed pilot never delegates merge approval, production deployment, vulnerability disclosure or spending approval to the scanner.

Separate the March product history from the September Cloud announcement

The wider Codex Security product predates the September Cloud announcement: OpenAI introduced it in research preview on 6 March 2026, with that announcement offering free usage for the following month. [Codex Security: now in research preview]

That historical offer is not evidence that a new October Cloud pilot is free.

The Cloud plugin is distinct from the local Codex Security plugin. The Cloud documentation describes connected GitHub repository scans in Codex Cloud; its local counterpart runs scans in a Codex task. [Codex Security Cloud FAQ]

This distinction matters when assembling the approval packet. Use the current Cloud setup and billing pages for the hosted experiment, not an old announcement’s promotional terms or a local-tool tutorial’s configuration examples. Record the surface being evaluated in the pilot’s first sentence. “Cloud plugin with connected GitHub repositories” is a more useful scope statement than a product-family name on its own.

We deliberately avoid a model ranking, an accuracy score and a claim that the hosted workflow is safer than a local one. None is needed to answer the immediate adoption question. The evidence packet should instead identify the repository, the permitted analysis environment, the people allowed to inspect results, and the action that each person may authorise. Keep purchasing expectations separate from technical curiosity: a useful demonstration is not yet a supported operating commitment.

Treat access and administration as separate approvals

For Enterprise and Edu workspaces, the help page requires both Codex Cloud and Codex Security access. Managing scan configurations additionally requires the relevant Codex Security administration permission; enabling use and enabling administration are separate decisions. [Codex Security]

Nominate an engineering owner for the pilot who can explain the selected repository and a security owner who can assess the finding evidence. Name an administrator separately. Where staffing requires one person to hold two roles, record that overlap rather than describing the process as independently reviewed. An approval document should identify who can stop the experiment while the usual administrator is unavailable.

Before connecting anything, obtain the repository owner’s approval for the assessment and the organisation’s approval for the processing arrangement. Confirm that the material may be shared with the proposed service under applicable contracts and policy. A developer’s ability to read source code should not be treated as permission to submit every surrounding dataset, incident record or customer attachment. Minimise those materials and use fabricated fixtures where they can answer the review question.

Ask the administrator to record the intended access, not merely confirm that the screen opens. List the selected repository, allowed participants and intended configuration changes. If an account appears to have broader reach than the charter allows, pause setup and reconcile the difference. Do not solve a missing repository by temporarily connecting the whole organisation and promising to narrow it later. An access problem is a reason to investigate, not a reason to bypass the boundary.

The October billing boundary belongs in the news lead

The current billing detail is decision-relevant enough to read before starting an evaluation. For source terminology, frequently asked questions (FAQ)A collection of recurring questions and concise answers about a subject. Open glossary entry refers to the Cloud questions page in the citations. Our spending recommendations below are approval rules for your team, not a claim that the product supplies a particular hard spending cap.

As of 11 October 2026, the Cloud billing guidance says repository scans and continuous scans set up after 1 October 2026 at 12:53 in the afternoon, Pacific Time, are billed at plan token rates, with eligible free scanning credits applied first. This usage is outside the plan’s included usage allowance. [Codex Security Cloud FAQ]

Continuous scans established before that cut-off remain free until 15 October 2026 under the documented transition. Continuing them afterwards requires enabling paid usage; otherwise they pause. Eligible accounts without that earlier continuous-scanning setup receive $500 in free scanning credits instead. The free balance is shared across repository and continuous scans and, within a workspace, across users. [Codex Security Cloud FAQ]

Do not translate the documented $500 credit amount into an assumed number of repositories or scans. Instead, ask the spending owner to confirm the account’s actual position: whether a transition applies, whether a free balance exists, which team shares it and whether paid continuation has been authorised. Write the answer beside the pilot start date. A colleague’s earlier trial may have had different terms, and a shared balance should not be treated as reserved money for your experiment.

A sensible pilot budget has two parts. Set a maximum authorised service spend under the organisation’s normal controls, and separately reserve reviewer time. Agree what happens when either is exhausted. If no suitable enforceable spending control has been verified, do not describe a written budget as a technical limit. Restrict the activity to an approved run, inspect its reported usage before another run, and stop if the owner cannot reconcile what would be charged.

Define a pilot that can answer one question

We recommend starting with one repository, adding a second only after the first review cycle, and reserving named reviewers before scanning. This is a deliberately bounded implementation of OpenAI’s recommendation to start with a small repository set and a dedicated reviewer group, not a product limit. [Codex Security]

Use an explicit learning objective: “Can our security reviewer and repository maintainer make a defensible disposition from the returned evidence, with an acceptable amount of follow-up work?” This is more useful than an objective such as finding as many vulnerabilities as possible. Quantity alone gives no decision rule for duplicate reports, weak evidence, urgent issues that need escalation, or a clean-looking result with unclear coverage.

The fictional Northwind Facilities engineering team will serve as the running example. Imagine that it owns a small maintenance-request service and a separate reporting component. These are planning examples, not customer experiences or reported product results. For a first pilot, the team chooses the maintenance-request repository because a maintainer can explain its access rules, a clean test fixture is available and the release process is understood. The reporting component remains outside the initial approval.

Reject repository choices that make the evaluation difficult to interpret. If nobody can explain the architecture, the pilot becomes an architecture-discovery exercise; if the only realistic test data contains sensitive records that have not been approved for processing, the pilot becomes a data-governance problem. If a service is under active incident response, its urgent remediation work should be governed by the incident lead rather than absorbed into a new-tool experiment.

Write a charter with observable exit conditions

The following fields are our suggested charter, not an official configuration format. Keep it brief enough that reviewers will read it before authorising a scan. Each field should have a named owner or an explicit unresolved status. A blank field is not a safe default. When a required decision cannot be made, the appropriate pilot state is “not ready”, not “approved subject to later clarification”.

  • Purpose: state the specific review decision the team wants to improve. Identify which existing process remains authoritative and explain how the experiment will avoid delaying ordinary security work.
  • Scope: name one authorised repository, the intended assessment target, any relevant history window and excluded materials. Record whether a second repository may be considered only after a fresh approval.
  • People: identify the repository maintainer, evidence reviewer, administrator, spending owner and incident escalation contact. Record cover for absence and who can suspend activity without waiting for a meeting.
  • Data handling: identify the permitted environment, fixture policy, evidence readers and retention decision. Describe how sensitive findings will be discussed without copying unnecessary source or personal information into broad channels.
  • Action boundary: allow analysis and review of proposed changes only. Require explicit human decisions for patch creation, sharing, merge, deployment, communication and paid continuation.
  • Stop conditions: list access mismatch, unexpected sensitive material, unexplained spend, unmanageable review backlog and unsafe validation activity. Assign the immediate containment action and the person who can approve resumption.
  • Exit evidence: require a completed review record for each investigated finding, an account of unreviewed results, a spending reconciliation and a written continue, narrow, pause or stop decision.

For the fictional team, an illustrative two-week evaluation window is enough to schedule preparation, one repository assessment, triage, a possible patch review and an end-of-pilot meeting. This duration is our planning suggestion, not a statement about scan speed or the product’s free-use period. The team should shorten or extend its calendar according to reviewer availability and keep any billing transition decision on its own dated track.

Choose a baseline without pretending to benchmark

Record how the selected repository is currently reviewed. Identify existing security checks, the last relevant human assessment and any already known issues that may appear again. The purpose is not to create a league table. It is to avoid counting a familiar defect as a new discovery or treating a difference in scope as a difference in quality. Keep known issues in a restricted reference list controlled by the security reviewer.

Before seeing results, define your disposition vocabulary. A useful starting set is: accepted for remediation, needs more evidence, duplicate of an existing issue, not applicable with a stated reason, and outside pilot scope. Do not force every uncertain finding into a true-or-false choice. A review process that permits an honest unresolved state is more informative than one that produces an apparently tidy dashboard by closing every difficult record.

Also define how reviewer effort will be recorded. Ask reviewers to distinguish time spent understanding the report from time spent rebuilding an environment, locating deployment context or writing a fix. These are different costs with different remedies. A weak report may need better evidence; an unfamiliar repository may need a better scope choice. Recording only one combined duration would hide that distinction and make the final decision less useful.

Agree what success will not mean

A quiet first run should not become a statement that the repository is secure, an impressive finding should not become a promise of the same outcome across the estate, and a proposed patch should not become a release instruction. These are pilot interpretation rules: they preserve the difference between an observation, a hypothesis, a tested change and a management decision. Put them in the opening briefing so they do not sound like excuses introduced after results arrive.

Do not ask the team to prove financial savings from a handful of examples. Instead, require a concrete account of usefulness: which decision became clearer, which uncertainty remained, whether the evidence could be independently checked, and which additional work was needed. If the answer is “the finding was plausible but the deployment assumption could not be verified”, retain that answer. It points to a context gap rather than justifying either promotion or rejection of the tool.

Finally, keep the pilot subordinate to qualified security judgement and organisational obligations. A novel workflow does not replace professional review, contractual duties, regulatory requirements or the organisation’s incident and disclosure procedures. If a credible serious issue appears, route it through the existing process immediately. The experiment can wait; the accountable security team must decide the appropriate response.

Read findings as evidence to inspect, not verdicts to accept

The reviewer’s task is to connect a reported concern to a specific piece of code, an explicit application assumption and an observable result. Begin with that chain rather than the severity label. For the fictional maintenance-request service, the relevant question might be whether a user is permitted to edit a particular request. An alarming description is not enough; the reviewer needs to understand the intended permission rule and the circumstances in which it might fail.

Conceptual illustration of security findings and evidence awaiting human interpretation
Conceptual illustration of security findings and evidence awaiting human interpretation. Original conceptual artwork, not a product screenshot or evidence of testing.

The Cloud documentation says findings include a description, source location, criticality, root cause and suggested remediation. Where verification steps are available, commands or tests run in the sandbox and their results are attached as evidence. [Codex Security Cloud FAQ]

Our suggested evidence packet keeps the original finding distinct from the reviewer’s interpretation. Preserve a private reference to the source record, then add the maintainer’s explanation of intended behaviour, the reviewer’s assessment and the current decision. Do not silently rewrite the original description into something more persuasive. If the reviewer changes the hypothesis, record the change and the reason so another person can reconstruct the discussion.

Ask four different questions of each finding

First, establish identity: which repository version and code location are under discussion? Second, establish the security rule: which actor should be prevented from performing which action? Third, establish the evidence: what was actually observed, under what conditions, and what remains an inference? Fourth, establish the response: what investigation or change has a person authorised? Keep these questions separate even when a generated explanation presents them together.

For example, the fictional report could concern a maintenance request being edited by the wrong team. The maintainer should identify the intended ownership rule without exposing real customer records. The reviewer should then check whether the supplied evidence addresses that rule or merely demonstrates an unusual input. A report about unusual behaviour may still deserve engineering attention, but it should not be upgraded into a security conclusion by the strength of its wording alone.

Auto-validation attempts to reproduce a suspected issue in an isolated container and records success or failure with logs, commands and related artefacts. If validation fails, the finding remains unvalidated; the record of what was attempted remains available for further investigation. [Codex Security Cloud FAQ]

For triage, distinguish failed reproduction from disproof. Our recommended “needs more evidence” state should record why the attempt did not answer the question: perhaps the expected fixture was missing, a dependency was unavailable, or the intended deployment assumption had not been established. These are illustrative possibilities, not observed defects in this product. The next step should target the missing premise rather than rerun the same investigation without a changed plan.

The converse also needs care. A reproduced behaviour in a controlled environment should be evaluated against the actual service context before the organisation assigns business impact. Ask which assumptions were supplied by the repository, which were supplied by a reviewer and which remain unknown. A security reviewer should approve the interpretation, while the service owner confirms operational context. Neither role should borrow the other’s authority merely to close the record faster.

Keep validation and build context separate

OpenAI says a compile step is not required to produce findings from repository and commit context. During auto-validation, the system may try to build the project if doing so helps reproduce an issue. [Codex Security Cloud FAQ]

A finding and a successfully built, representative test environment are therefore separate things to verify.

Before accepting a validation record, check whether the environment represents the decision you need to make. For an access-control concern, a fixture with the right actor relationships may be more useful than a large collection of unrelated application data. For a parsing concern, the reviewer may need a narrowly scoped test and a clear explanation of expected handling. Prefer the smallest authorised fixture that answers the question, and have the maintainer approve its relevance.

Do not obtain production records simply because they would make a reproduction feel more realistic. Ask whether synthetic data can preserve the property being tested. If it cannot, stop and seek approval for a different assessment route. The pilot should not become an informal exception to data handling policy. Record when realism was limited and how that limitation affects confidence, rather than allowing the report to imply a production-equivalent test.

Keep environment changes in the evidence history. If a reviewer adds a dependency, changes a fixture or corrects an assumption, note what changed before comparing the new result with the previous one. Otherwise, the team may attribute an improvement to the scanner when it actually came from better context. This is a recommended review practice, not a claim that the Cloud interface provides a complete experimental versioning system.

Use the threat model as reviewable context

For a monitored repository, Codex Security drafts the threat model from code and uses it to guide future commit scans and finding prioritisation. The editing instructions put Threat model under Project context in Monitoring settings, and explicitly say changes apply to future scans. [Improving the threat model]

In your own review record, treat the threat model as a set of propositions the application owner can confirm or challenge. Write short statements about entry points, permitted actors, sensitive operations and trust assumptions. Avoid a long architecture narrative that does not change how a finding is assessed. The best context for this pilot is not the most comprehensive document; it is the smallest accurate account of the boundaries relevant to the selected repository.

A useful correction for the fictional team might be: “A maintenance coordinator may change requests assigned to that coordinator’s team; ordinary requesters may edit only their own drafts.” That is an illustrative business rule, not a product prompt or a real policy. Before using such a statement, the actual owner must confirm it against the application’s intended behaviour. Do not let generated source interpretation become the definition of the organisation’s access policy.

When the context changes, note the date and the reason in the pilot record. Keep earlier findings associated with the assumptions under which they were produced. Our recommendation is to make a fresh review decision about affected records, rather than assuming a context edit retroactively settles them. Ask whether the new statement changes the likelihood of the concern, its potential impact, or simply the explanation reviewers need.

Triage without losing uncertainty

The setup guide places New, Triaged and In Progress findings in Open findings, with search and Filters for narrowing the list. Selecting an issue opens affected code, validation evidence and remediation guidance. [Codex Security Cloud setup]

Use the product’s status labels for navigation, but keep the pilot’s assessment vocabulary explicit in your review notes. “Triaged” should not be treated in your management report as synonymous with “confirmed”, “fixed” or “safe”. Define the relationship between the status visible in the tool and the decision recorded by the reviewer. If the same word means different things to security and engineering, agree on the meaning before preparing a summary.

A false positiveA reported positive finding that a later check shows is not a genuine match; in code review this may be a flagged rule violation that the labelled fixture does not actually contain. Open glossary entry is one possible disposition, but the label should follow an explanation. For example, a reviewer may determine that a supposed permission boundary does not exist in the authorised application design. Record the supporting context and who confirmed it. Do not use the label merely because reproduction was inconvenient, the issue would be expensive to fix, or a generated patch failed to satisfy the maintainer.

For duplicate handling, our recommendation is to group reports by the underlying security hypothesis and affected component, not by title alone. Keep references to each source record and document why they are considered the same issue. If two reports require different fixes or rely on different trust assumptions, preserve them separately until a reviewer resolves the relationship. This manual discipline is useful regardless of any automatic grouping visible during the pilot.

Protect the evidence packet as security-sensitive working material. Restrict readers to those who need the details, minimise copied code and avoid broadcasting reproduction material in routine project updates. A management summary should state the decision and owner without carrying every investigative detail. Where disclosure to a supplier or maintainer is needed, the authorised security contact should decide the channel, timing and content under the existing disclosure process.

What is not yet documented in the reviewed sources

This section records the limits of the source pack reviewed on 11 October, not an assertion that an undocumented capability cannot exist. A missing answer should become a question for the administrator or product support when it is material to approval. Do not fill gaps by borrowing assumptions from another Codex surface, extrapolating a demonstration or treating the word “cloud” as a complete security and operations specification.

The current Cloud questions page calls the tool language-agnostic but qualifies performance by the model’s reasoning ability for the repository’s language and framework. It also says scan time varies with repository size and validation work; neither statement supplies a guaranteed completion time or equal performance across languages. [Codex Security Cloud FAQ]

For this pilot, we have no source-backed basis for promising a particular duration on the selected repository, equal detection quality across all languages, complete vulnerability coverage or a guaranteed number of actionable findings. The approval should say that these are evaluation questions. If the business needs a contractual completion deadline or a specific coverage guarantee, obtain that assurance separately before making the workflow dependent on it.

The documentation describes each analysis and validation job as an ephemeral Codex container with session-scoped tools. It says artefacts are extracted for review before container teardown. [Codex Security Cloud FAQ]

Container teardown must therefore not be read as a statement that all review evidence has been deleted.

We did not find, in the setup and questions pages we reviewed, the detailed evidence-retention period, deletion verification process or region-by-region eligibility table that many organisations need for approval. We therefore do not assert a universal retention duration, a particular processing location or worldwide availability without qualification. Ask the responsible privacy and security reviewers to confirm the applicable contractual and account-specific terms. If those terms are a condition of processing, do not start while they remain unresolved.

Similarly, this article makes no promise of private-network reachability, a particular repository-hosting deployment variant, automatic access revocation across all artefacts, or a comprehensive audit export. These are requirements to investigate if your organisation needs them. Write each requirement as an acceptance question with an owner and evidence source. A favourable answer about one of them should not be used as a substitute for the others.

The DevDay announcement uses “on demand or on a schedule”, while the setup guide names One-Time Scan and Continuous Scanning. Continuous Scanning is documented against the repository’s default branch. We recommend designing the pilot around those named controls rather than assuming a configurable scheduling interval or coverage of every branch. [DevDay 2026 Recap; Codex Security Cloud setup]

Use the same restraint for coverage language. The pilot’s scope statement should distinguish a repository assessment from ongoing review of changes. Do not present a scan of one selected target as a guarantee about unselected branches, linked services, deployed infrastructure or repositories outside the approval. If a finding depends on an external component, record the dependency and route it to the responsible owner rather than expanding the assessment without permission.

Do not assume an included entitlement travels elsewhere

Daybreak Blue access included with Codex Security Cloud is scoped to the Cloud product. The help page says it does not provide Daybreak Blue access in other Codex Security products or through the application programming interface. [Codex Security]

For an engineering lead, the operational consequence is to avoid planning a follow-on integration around access that has not been approved on that integration’s own terms. Evaluate the hosted workflow first. If a later proposal requires a different product surface, create a new access and security review rather than describing it as a routine extension of the pilot. This article does not supply a model entitlement map or an integration approval.

Keep a small documentation-gap register

For each unresolved requirement, record the question, why it matters, who can answer it and whether it blocks the next action. For example, uncertainty about evidence retention may block uploading a sensitive repository but not a planning meeting using fictional material. Uncertainty about billing may block a scan but not review of the published setup instructions. This separation prevents an entire discussion from stalling while still preserving the necessary stop boundary.

Use precise resolution criteria. “Security is comfortable” is not enough for a requirement about who can read findings. A useful resolution identifies the applicable policy, the reviewed setting or contractual answer, the approver and the allowed scope. If a setting is not documented or cannot be verified, preserve that uncertainty. Do not turn a support discussion into a public product claim unless there is suitable attributable evidence for publication.

At the end of this stage, the pilot should have a concise evidence standard, an agreed disposition vocabulary and a list of unresolved conditions. That preparation is the point of the documentation review. The team is ready to run a bounded assessment only when it knows what it will inspect, which conclusions it is not entitled to draw and which person can authorise the next step.

Run one controlled review cycle before enabling a second repository

The sequence below is a proposed pilot procedure built around the documented controls. It is not a claim that we performed the steps or observed a successful result. Have the administrator and repository owner approve the charter first, then work through the sequence with a security reviewer available. Keep ordinary repository protections and change approvals in place throughout the evaluation.

1. Confirm the connection against the approved scope

The setup sequence is to open Plugins, find and enable Codex Security Cloud, then open Security Cloud. Inside the plugin, Scan leads to repository selection; Connect GitHub is offered when needed, and a missing repository should prompt a check of the connection and permissions. [Codex Security Cloud setup]

Before granting repository access, compare the proposed selection with the charter. Confirm the organisation, repository owner and purpose in your own approval record. If the expected selection is unavailable, inspect the permission issue with the administrator rather than borrowing another person’s account. A workaround that changes identity or scope changes the experiment and should receive a new decision.

After setup, ask a second authorised participant to inspect the intended pilot access and explain what they are permitted to do. This is our recommended peer check, not a claim about a particular built-in test mode. The aim is to catch misunderstandings before any findings become sensitive shared material. Do not grant the second participant additional privileges merely to make the check convenient.

Keep screenshots or administrative evidence only where policy permits, and redact unrelated repository names or personal information before using them in a wider approval discussion. The private pilot record should preserve enough detail to explain the decision without becoming a second uncontrolled copy of the organisation’s source inventory. If the review exposes unexpected access, stop and resolve that exposure before evaluating scan quality.

2. Start with a single assessment and record its target

In New Scan, the documented Auto environment option creates an environment when the scan starts; Customize selects an existing one. For a single repository assessment, the setup guide says to choose One-Time Scan under Scan Method, select Create, then follow progress and artefacts in Scans. [Codex Security Cloud setup]

Have the maintainer review the proposed environment before starting. Identify the dependencies and fixtures necessary to interpret a finding, and the information that must not be included. Do not add production credentials as an expedient response to a build problem. If a required dependency cannot be supplied within the approved boundary, narrow the evaluation or record that this repository is not ready for the pilot.

Record the target version as precisely as the workflow and your own repository records allow. Note when the scan was requested, who requested it and which settings were selected. If the repository changes during analysis, do not assume the resulting evidence refers to the newest code. Reconcile the source reference before assigning a fix. Where the reference cannot be established, label that gap rather than attaching the result to a convenient current version.

Use the initial assessment to test the review process, not to consume all available capacity. Reserve time to read incomplete or unclear results, examine validation evidence and ask the maintainer about context. If the team starts additional runs before understanding the first one, the pilot may become a queue-management problem. Our recommendation is to authorise each early repeat deliberately and state what new question the repeat is intended to answer.

3. Assign a human disposition to each investigated issue

Hold a short joint triage session after the first results are ready for inspection. The security reviewer should explain the suspected boundary failure; the maintainer should explain intended behaviour and relevant deployment assumptions. Compare their interpretations before selecting a disposition. If they disagree, preserve both views and identify the evidence needed to resolve the disagreement rather than averaging them into an ambiguous medium-priority label.

For the fictional maintenance-request service, imagine a concern about cross-team edits. The review should ask whether team membership is checked at the relevant decision point, what the test fixture represents and whether any separate control changes the interpretation. These are questions for an authorised review, not claims that the scanner found this issue. Do not construct or run an exploit against a live service as part of this article’s pilot.

Assign an owner and next action even to unresolved findings. “Needs more evidence” should say who will obtain that evidence and when the record will be reconsidered. For a duplicate, identify the existing issue and explain why it is the same concern. For an out-of-scope result, identify the receiving owner without copying more sensitive detail than necessary. The pilot remains useful when it exposes a gap, provided the gap is not disguised as completion.

4. Treat a proposed patch as a separate change request

Conceptual illustration of a human-reviewed patch and monitoring decision
Conceptual illustration of a human-reviewed patch and monitoring decision. Original conceptual artwork, not a product screenshot or evidence of testing.

When a finding offers Fix with Codex, the documented action generates a proposed patch. The guide instructs the reviewer to inspect that patch before selecting Create draft pull request. The availability of this action is conditional: not every finding is promised a patch. [Codex Security Cloud setup; Codex Security Cloud FAQ]

The Cloud questions page says patches are not auto-applied. It describes a generated diff, patch file or suggested change for maintainers to inspect before applying, rather than a direct modification of the existing pull-request branch. [Codex Security Cloud FAQ]

Approve patch generation only after the team understands the issue it intends to address. Our suggested change boundary is one confirmed or explicitly investigated concern with a named maintainer. Ask the maintainer to reject unrelated cleanup, broad refactoring and dependency changes that are not necessary for the proposed remedy. If the narrow fix reveals a larger design question, move that question to the normal architecture or security process rather than expanding the pilot silently.

Read the change against the security rule, not only against the report’s wording. A patch can make a demonstration stop working without preserving the application’s intended behaviour. The reviewer should therefore ask what remains allowed, what becomes disallowed and whether the change shifts responsibility to another component. Record any new assumption introduced by the fix. The acceptance question is whether the authorised behaviour remains correct while the identified concern is addressed.

Require separate tests for the restricted action and the legitimate action. For the fictional service, a review plan might check that an authorised coordinator can still update an assigned request and that an unrelated actor cannot make the same change. These are illustrative test objectives, not executable instructions or reported outcomes. The maintainer should choose the actual fixtures and assertions under the repository’s established test process.

Also examine failure behaviour. Consider how the proposed change should respond to missing information, stale state and invalid input within the approved test environment. Avoid treating every failure as acceptable merely because access was denied. An unnecessarily broad denial may break a legitimate workflow. The change owner should describe the intended user-facing outcome and the recovery path, while the security reviewer confirms that the boundary remains appropriate.

5. Make draft review ownership explicit

GitHub documents that a draft pull request cannot be merged and does not automatically request code-owner reviews. Marking it ready for review requests reviews from code owners. [Pull requests]

Draft creation should therefore be followed by an explicit human assignment and review decision, not treated as proof that the right reviewer has been notified.

Assign the reviewing maintainer deliberately and tell that person what is ready for assessment. A draft should have a clear reason for remaining a draft: incomplete tests, unresolved assumptions, pending security review or another stated condition. Avoid using draft status as a substitute for an owner. A record that cannot merge may still contain sensitive details and may still leave a real issue unresolved.

Before any change is marked ready, require the patch author or responsible maintainer to prepare a concise review note. It should identify the finding, the intended security rule, the scope of the proposed change, tests completed by authorised people or systems, and unresolved limitations. Distinguish observed outcomes from tests merely requested. An unexecuted test plan must remain a plan; it should never appear in the review note as a successful check.

The following is a copy-ready human handoff template. It is our editorial template, not an official product schema. Keep it in the organisation’s approved private review system and remove fields that would duplicate sensitive evidence unnecessarily. A person should complete and verify it; the presence of a filled field does not prove its accuracy.

Repository review handoff
Authorised scope: name the approved repository and assessment target.
Finding reference: link within the approved private evidence system.
Security rule: state the intended actor, action and boundary.
Evidence assessment: separate observation, inference and unknowns.
Proposed change: describe the smallest intended correction.
Tests actually completed: name the reviewer, target and observed result.
Tests still required: state the unresolved question, not an assumed outcome.
Data and rights check: confirm authorised use and minimise copied material.
Review owner: name the maintainer accountable for the decision.
Decision: remain draft, request changes, or seek normal approval.
Authority boundary: no automatic merge, deployment, disclosure or spending.

Do not attach an unredacted evidence bundle to a broadly visible change description just to make the handoff look complete. Keep a minimal summary with an approved private reference when policy allows that pattern. The security owner should decide whether a normal pull request is an appropriate place for the detail at all. Particularly sensitive issues may require the organisation’s established restricted remediation route.

6. Authorise monitoring only after the first evidence review

For continuous monitoring, the setup guide offers Scan commit history from, optional threat-model scoping guidance and a selected cloud environment. Longer history windows add context but make the initial scan take longer. Monitoring settings allow changes to environment and history, with Save applying changes. [Codex Security Cloud setup]

Make monitoring a second decision, not the assumed continuation of a successful first scan. The approval should identify who watches the results, how often the team reviews the queue and what happens during absence. Choose the history window for a stated question rather than maximising it by default. If historical context is relevant, explain why; if the immediate purpose is reviewing new changes, keep that purpose visible in the charter.

Before switching to ongoing activity, estimate review capacity in records per review session without inventing a product throughput figure. The team can choose its own workload ceiling, such as accepting only as much new triage work as the named reviewers can inspect before their next scheduled session. This is an internal staffing rule. It should trigger a pause when capacity is exhausted rather than encourage hurried acceptance of weak evidence.

Keep the second repository outside scope until the first cycle has an intelligible result. If the first repository produces a useful patch but requires extensive manual environment repair, decide whether that repair is reusable or specific to the chosen project. If the evidence is mostly unresolved, improve context before adding more inputs. Expansion should follow a reasoned assessment of the constraint, not the availability of another repository selector.

7. Practise the stop decision before you need it

To pause commit monitoring, the Cloud instructions say to open Repositories, select the repository, open Monitoring settings, set monitoring to Paused and select Save. [Codex Security Cloud FAQ]

This documents a monitoring control, not a claim about deleting historical findings or cancelling every job already in progress.

At the pilot briefing, walk through a tabletop stop exercise. Ask the administrator to explain the documented pause sequence without changing production settings. Ask the spending owner what they would do if usage could not be reconciled. Ask the security reviewer who receives an unexpected sensitive finding. The purpose is to remove ambiguity about authority, not to claim that every technical failure mode has been tested.

If an incident occurs during the actual pilot, record the last confirmed state, stop initiating additional activity and use the organisation’s approved containment process. Do not assume a pause removes all prior access or erases results. Have the administrator verify each required containment action separately, including any connection or permission change the incident owner requests. Preserve necessary evidence under policy while avoiding unnecessary distribution.

Restart only with a written explanation of what changed. That explanation should identify the root of the pilot failure where known, the evidence supporting the correction and any remaining uncertainty. A repeated scan without a revised hypothesis or corrected boundary is not a recovery plan. It simply produces another result that reviewers may be unable to interpret.

Decide whether to continue using evidence, workload and control

The end-of-pilot decision should reconcile three records: what the assessment produced, what reviewers could establish, and what the organisation authorised or spent. Keep these records connected but do not collapse them into one success score. A useful finding does not excuse an access error, a well-controlled run does not prove detection quality, and low expenditure does not make an unreviewed backlog an acceptable security outcome.

Reconcile usage before approving another cycle

The Cloud questions page says that opening a scan in Scans shows its token usage and cost; hovering over the token count separates input, cached input and output. Cached input is included within input, not added to it. The displayed cost is before free credits or exemptions, with covered amounts also shown. [Codex Security Cloud FAQ]

Have the spending owner reconcile the pilot’s activity with the account’s actual billing position. Record which runs belong to this pilot, which allowances or credits were applied and which other workspace activity may share the balance. Avoid publishing a cost-per-repository estimate from one run. If the cost display and the account record do not appear to agree, seek clarification before initiating another assessment rather than assuming the smaller number is the authoritative charge.

The decision record should distinguish authorised spend, displayed usage cost, credit coverage and the amount actually charged or expected under the account’s terms. These are accounting categories for the review, not a new billing formula. Preserve the evidence the spending owner used and the date of reconciliation. If a figure is provisional, label it provisional and explain what must happen before it can support a renewal decision.

For continuous scans covered by the transition, the documented choice is Keep scans running to enable paid usage. Opting out or not enabling it pauses those scans at the end of the free period; Re-enable scans can enable paid usage afterwards. A workspace owner must make the choice if the user lacks permission. [Codex Security Cloud FAQ]

If your team is reading this on 11 October, the documented 15 October transition deserves its own calendar decision if it applies to their existing scans. The action should belong to the authorised workspace or spending owner, not to whichever engineer sees a continuation prompt first. Our recommendation is to decide before the boundary whether the pilot should continue, pause for review or end. Do not approve paid continuation merely to preserve a demonstration for an upcoming meeting.

Measure review value without inventing a benchmark

Use the evidence standard agreed before the run. Count investigated findings by disposition, and keep unreviewed findings visible as a separate category. A record that is awaiting a reviewer is not the same as one found to be unconvincing. A record accepted for remediation is not necessarily new to the organisation. The report should make those distinctions clear enough that a manager cannot mistake queue movement for risk reduction.

If you report a confirmation proportion, state the denominator precisely: for example, findings reviewed to a final disposition during this pilot, with unresolved records listed separately. Do not call that proportion a universal accuracy rate. The sample was selected for a local decision and may reflect reviewer availability, repository familiarity and the scope of the assessment. Keep the original counts alongside any percentage so a small denominator remains obvious.

Track evidence completeness independently of disposition. Ask whether a reviewer could locate the relevant code, identify the target version, understand the application assumption and determine what validation attempted. A rejected concern can still arrive with a useful evidence trail. An accepted concern can still require substantial manual reconstruction. Those differences help the team decide whether the next investment should be in better context, a different repository choice or improved review procedure.

Assess patch usefulness separately again. Record whether a proposal was offered, whether the maintainer chose to investigate it, whether the change stayed within the approved scope and what additional work was required. Do not count a draft as a merged fix or a merged fix as a deployed remediation. Each step requires its own evidence and accountable owner under the organisation’s normal process.

OpenAI explicitly says Codex Security does not replace manual security review, code-level validation, exploitability checks or human threat assessment. We recommend judging the pilot as an additional evidence source, not as approval to retire the existing review process. [Codex Security Cloud FAQ]

In the final meeting, ask reviewers to describe one specific decision the evidence helped them make and one uncertainty it did not resolve. If there was no useful decision, record that plainly. If the main value was clarifying the application’s threat assumptions, say so rather than describing the pilot as successful automated remediation. An honest narrow benefit is more actionable than a broad claim that cannot be traced to the review records.

Apply stop conditions consistently

Use the following operating responses as recommendations for the pilot team. They are not descriptions of automatic product safeguards. The owner named in the charter should authorise containment and any subsequent restart. Where an event becomes an organisational incident, leave the experiment’s normal cadence and follow the incident lead’s instructions.

  • Unexpected repository access: stop initiating scans, preserve the minimum record of the discrepancy and ask the administrator to reconcile access with the charter. Resume only after the repository owner confirms the corrected scope.
  • Unapproved sensitive material: restrict further handling, notify the appropriate security or privacy owner and follow the established containment and retention process. Do not circulate the material to demonstrate the problem.
  • Unexplained usage or funding: stop new paid activity and ask the spending owner to reconcile the account. Treat an unresolved credit or billing assumption as a decision blocker, not as a reason to continue until a charge appears.
  • Unsafe validation proposal: do not run it against production or a third party. Route the assessment to a qualified reviewer who can define an authorised, isolated alternative or decline further investigation.
  • Review backlog exceeds capacity: pause expansion and consider pausing monitoring under the approved process. Prioritise existing records with the security owner rather than lowering the evidence standard to clear the queue.
  • Patch exceeds the agreed concern: return it for revision or move the broader change into the normal engineering process. Preserve the original finding while separating remediation from opportunistic refactoring.
  • Credible urgent vulnerability: escalate through the existing incident or vulnerability-management route. The pilot schedule must not delay a qualified human response or dictate disclosure timing.

Choose a stop threshold that can actually be observed. “Stop if anything seems wrong” gives neither the operator nor the sponsor a usable instruction. Prefer a concrete mismatch, an unanswered approval question or an agreed workload boundary. At the same time, allow any participant to raise a concern without needing to prove an incident. The response can begin with a precautionary pause and a proportionate review.

Keep the record of a stopped run neutral. Describe what happened, what was known at the time, who decided to stop and which evidence remains available. Avoid labelling the scanner unreliable merely because the environment was unsuitable, or labelling the team unprepared because the documentation did not answer a required question. The purpose is to identify the limiting condition and decide whether it can be resolved within the approved scope.

Use a small weekly review rhythm

For a continuing pilot, begin each review session with scope and access, not the finding count. Confirm that the selected repositories, named reviewers and spending authority are unchanged. Then inspect newly available evidence, unresolved records and proposed patches. End by assigning owners and deciding whether the next period should have the same scope, a narrower scope or no further scanning.

Reserve a separate moment for threat-context changes. Ask the maintainer whether application behaviour, ownership rules or deployment assumptions have changed since the previous session. Where a change matters, record it before interpreting new results. Do not rewrite the history of earlier assessments. A later understanding of the system may justify revisiting a decision, but the review record should preserve why that earlier decision was made.

Keep a short management summary that does not expose sensitive detail. Include the current pilot state, the decisions made, the unresolved blockers, the next human owner and the spending position. Avoid a list of dramatic vulnerability titles without context. For stakeholders who do not review code, the useful information is what the team knows, what it does not know and which action has been authorised.

Choose continue, narrow, pause or stop

OpenAI’s best-practice guidance recommends normal review for generated patch pull requests. Our proposed expansion gate therefore requires a named maintainer’s patch decision and a security reviewer’s evidence assessment before increasing repository scope; the gate is our operating recommendation, not a built-in approval feature. [Codex Security]

Continue at the same scope when reviewers can trace the evidence, the workload is manageable and the next cycle has a specific learning objective. This choice does not require expanding to another repository. Repeating a well-defined process can be the appropriate next step when the team is still learning how to interpret results or improve context. State what the repetition should clarify and when the decision will be revisited.

Narrow the scope when the first repository or history window introduces more uncertainty than the team can resolve. A smaller approved target, a more focused question or better fixtures may produce a more interpretable evaluation. Narrowing is not a failure if it makes the evidence useful. Record what was excluded so that later stakeholders do not mistake a more manageable experiment for unchanged coverage.

Pause when a remediable prerequisite is missing: an approval, a suitable environment, a named reviewer or a reconciled budget. Give the pause an owner and a resumption condition. Do not leave it as an indefinite state that nobody checks while assuming monitoring continues. The decision record should specify what has stopped, what evidence remains to be reviewed and which controls the administrator has actually confirmed.

Stop when the workflow cannot meet the organisation’s requirements within reasonable effort or when the benefit is not sufficient for the review burden. Close outstanding records responsibly, preserve necessary evidence under policy and review access removal through the normal administrative process. Ending the pilot should not erase a credible finding or leave a proposed patch without an owner. The tool decision and the vulnerability decision are separate obligations.

Write an exit memo that another lead can challenge

The exit memo should be short enough to read before a meeting but specific enough to audit afterwards. Begin with the authorised question and scope, then explain the evidence actually reviewed. Describe the most useful result without exaggeration, the principal unresolved limitation and the operational cost. Finish with the chosen action and the person responsible for carrying it out. Attach private references rather than copying the entire evidence collection.

A useful challenge test is to ask another engineering lead to identify the strongest reason not to expand. If the memo contains only favourable examples, it probably does not support a reliable decision. Invite the security reviewer to identify any finding whose disposition depends on an unconfirmed assumption, and the spending owner to identify any unresolved charge. These are review questions, not reasons to create an elaborate scoring model.

For the fictional Northwind team, several outcomes would be legitimate: continuing on the maintenance-request repository, postponing the reporting component until its data boundary is clearer, or ending the experiment after deciding that evidence review costs too much attention. This article does not assign an outcome to that fictional pilot. Its purpose is to give the real team a decision structure without fabricating a success story.

A preview is worth evaluating only within a clear boundary

The practical recommendation is to keep the first experiment small, make evidence review explicit and preserve human authority over every consequential action. Read the current billing terms before a scan, distinguish validation from business impact, and treat every patch as a proposal. Expand only when the named owners can explain both the value and the unresolved limits. A well-documented decision to pause is a better outcome than an unreviewed promise of security.

Explore the Prompt Library for ChatGPT, Claude & Codex

Subscribe to access the curated Notion Prompt Library, with practical prompts organized for coding, research, content creation, and business workflows.

Access the Prompt Library →

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this