Before You Trust Codex Code Review: A Reproducible PR Benchmark for Rule Signal, Reviewer Time, and False-Positive Noise

A balanced tabletop scale holding a small cluster of bright markers on one side and scattered dull markers on the other, with an empty measuring grid beneath.

Protocol status: this is an unrun benchmark protocol. This site has not executed it, observed its outputs or measured Codex Code Review. It therefore provides no product rating, recall or precision result, reviewer-time saving, cost saving, testimonial or rollout recommendation. Every threshold and sample record below is a proposed method to be fixed before a run, not a reported outcome.

A balanced tabletop scale holding a small cluster of bright markers on one side and scattered dull markers on the other, with an empty measuring grid beneath.
The pilot is designed to weigh consequential review signal against the human cost of noisy comments—not to assume that more comments mean better review.

Evidence checkpoints

Documented point: As reviewed on 2 October 2026, OpenAI documents @codex review, automatic reviews for connected repositories, GitHub comments focused on its highest-severity P0 (priority-zero) and P1 (priority-one) categories, scoped AGENTS.md guidance and the continuing need for tests and human approvals. This establishes the Code Review surface and setup behaviour, not recall, precision, security coverage, or time savings in a particular repository. [official source 1]

Documented point: In an OpenAI internal suite described on 20 July 2026, rule-guided variants recovered 98% of required custom findings versus 58.3% in a baseline control; those figures describe that suite, not this unrun benchmark. These results do not establish performance in another repository. [official source 1]

Documented point: As reviewed on 2 October 2026, OpenAI recommends explicit objectives, representative datasets, defined metrics, logged results and automated scoring calibrated against human judgements; its Evals platform was scheduled to become read-only on 31 October 2026 and to shut down on 30 November 2026. The proposed method uses durable repository and continuous-integration artefacts; the retiring Evals platform is not required. [official source 1]

Documented point: As reviewed on 2 October 2026, OpenAI distinguishes exploratory trace grading from repeatable datasets and evaluation runs for comparing changes over time. This is evaluation guidance, not evidence that Codex Code Review exposes identical trace/evaluation functionality or that a proposed score has been obtained. [official source 1]

Documented point: As reviewed on 2 October 2026, Codex access is included across ChatGPT plans, while Codex Cloud eligibility varies by plan, rollout and workspace settings; usage also varies by plan and task. Check current plan, workspace and regional eligibility; costs vary by usage and product surface. [official source 1]

Documented point: As reviewed on 2 October 2026, OpenAI distinguishes Chat, Work and Codex: ChatGPT-account Codex sign-in follows plan usage and billing, whereas a separately supplied application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry key follows API pricing. Treat ChatGPT-plan and separately billed API usage as distinct cost categories; do not use API-token prices to estimate a GitHub review pilot without evidence of its selected product surface. [official source 1]

The decision is not whether a comment appeared

This is an unrun study design, not a tested Codex Code Review benchmark: no experiment was conducted or measured result obtained here. Use a sanitised fixture with no production secrets; human approval is mandatory before accepting a finding, changing review rules or making a release decision.

A GitHub review comment proves only that a comment was posted. It does not show that the comment identified the intended defect, distinguished an unsafe change from a safe counterexample, preserved detection of unrelated serious defects, gave a reviewer enough evidence to act, or justified the time spent triaging it. Even a technically correct comment can make a review process worse if it duplicates another finding, assigns the wrong severity, points at the wrong location or recommends an unsafe remedy.

The rollout question should therefore be framed as a comparative decision: under fixed and recorded conditions, do a few repository-specific AGENTS.md rules improve detection of consequential, pre-labelled rule violations without unacceptable losses in ordinary-defect detection, false-positive noise, actionability or human triage time? The comparison matters. A collection of impressive-looking comments from the rule-guided condition alone cannot establish that scoped rules caused an improvement. The same pull request (PR)A proposed set of repository changes submitted for review before integration. Open glossary entry cases must also be reviewed under defined controls.

The decision owner should be named before the first run. For this protocol, that person is normally the developer-experience lead, staff engineer or repository owner who can accept the reviewer-attention cost and decide whether the rules remain advisory. The owner is not merely the person operating the benchmark. They must approve the defect-severity policy, false-positive tolerance, time budget, exclusion policy and stop conditions. Security, privacy, financial, employment, government and other consequential decisions require appropriately authorised human review; neither a Codex comment nor a benchmark score should decide them automatically.

Define the decision product before defining the metrics

The decision product is a bounded recommendation about one connected GitHub repository and one recorded review workflow. It can support “retain these scoped rules for a larger advisory pilot”, “rewrite and rerun”, or “do not expand this rule set”. It cannot support “Codex is accurate”, “artificial intelligence (AI)Computer systems designed to perform tasks that normally require human intelligence, such as understanding language, recognising patterns, or making predictions. Open glossary entry review saves time”, “this plan is cheaper than an application programming interface workflow”, or “the repository is secure”. Those broader conclusions require different evidence.

Use the following operational definitions in the pre-registration. Changing them after reading condition C’s output would make the comparison vulnerable to outcome-driven relabelling.

Repository-specific rule signal
The change in correctly recovered, pre-labelled repository-rule defects between conditions, measured at defect level rather than by counting comments. A finding contributes signal only when it identifies the labelled defect and explains its consequence or safe path. A generic style observation on the same lines does not count as recovery.
False-positive burden
The human-facing load created by comments that do not identify a valid defect, especially comments on deliberately safe counterexamples. Report both false-positive comments per clean PR and the reviewer minutes required to dismiss or investigate them. One noisy PR with five comments is not equivalent to five clean PRs with one comment each, so retain case-level distributions.
Ordinary P0/P1 defect retention
The extent to which rule-guided review continues to identify pre-labelled serious defects unrelated to the custom rules. Here, P0 and P1 must follow the repository’s written severity policy. OpenAI documents its GitHub comments as focused on P0/P1 issues; low-severity style preferences should not be relabelled as P1 merely to increase the denominator.
Actionability
Whether a finding gives a human enough grounded information to make a safe next decision. The rubric should require an affected location, a real consequence, adequate evidence, a safe remediation path and an appropriate severity. Actionability is distinct from correctness: a correct but vague warning can recover a defect while failing the actionability gate.
Triage time
Active human time spent reading, checking, classifying and deciding what to do with the review output. It excludes unattended waiting for a review to appear and excludes implementation time unless the benchmark explicitly adds a separate measure. Review latency and human triage time must be reported separately.
Usage proxy
Any visible plan credit, allowance or usage indication captured for the run, alongside the account and workspace configuration. It is a proxy rather than a monetary cost. OpenAI’s Help Centre, updated 2 October 2026, distinguishes ChatGPT-account Codex usage from use with an organisation’s own API key; a ChatGPT-plan run must not be converted into API-token pricing.

Decision rule: if the owner cannot identify a consequential rule outcome, a tolerable clean-PR noise level and an available human-review budget, do not run the pilot yet. A benchmark with no pre-agreed decision boundary can produce charts, but it cannot provide a defensible rollout decision.

Hold the product surface constant

OpenAI’s connected-GitHub documentation, checked 2 October 2026, describes a manual @codex review trigger, automatic reviews for connected repositories, comments focused on P0/P1 issues (its labels for high-priority findings), and applicable root or nested AGENTS.md guidance. It also states that this guidance does not replace tests, branch protections or required approvals. These facts define the benchmark surface; they do not establish its performance in a particular repository.

The protocol may use either manual or automatic review, but not silently mix them. Select one surface for the primary comparison and record it on every run. If both are studied, treat them as separate cohorts because trigger timing and repository settings may differ. A missing review is a run outcome, not permission to substitute a result from another Codex surface.

Product boundaries for this benchmark
Surface What distinguishes it Treatment in this protocol Decision rule
Connected-GitHub Codex Code Review Reviews a GitHub PR through the documented manual or automatic-review route and can apply relevant repository guidance. This is the only evaluated product surface. Preserve the PR, diff, review, timestamps, exact commits and applicable rule-file revision. Include a run only when the pre-registered trigger and connected repository are verifiable.
Consumer ChatGPT An ordinary chat experience is not the connected-GitHub PR review workflow. Exclude chats, pasted diffs and conversational code critiques from the benchmark dataset. Do not use a useful chat answer as replacement evidence for a missing GitHub review.
ChatGPT Work OpenAI’s Help Centre distinguishes Chat, Work and Codex as different experiences. Exclude Work tasks and outputs, even if they concern the same repository. If work was performed outside the PR review surface, label it separately and do not merge its scores.
Codex Cloud A remote execution surface in which tasks use configured environments and workspaces; it is not itself evidence about GitHub review behaviour. Exclude cloud task outputs, generated patches and cloud implementation results. Run a separate benchmark if the decision concerns implementation tasks rather than review comments.
Codex Security A security-oriented surface with a different purpose from ordinary PR review. Exclude scans, vulnerability-coverage claims and security certification language. Any security-relevant comment must still receive qualified human security review; it cannot validate repository security.
Local Codex Local execution has its own environment, sandbox, approval and network conditions. Do not substitute command-line or local-agent output for the GitHub review artefact. If local tooling is used only to collect or analyse public artefacts, record that separately from the reviewed product.
API workflow An API implementation has separately selected models, prompts, orchestration and API billing. Exclude API calls and model comparisons from the primary benchmark. Do not infer an unexposed model identity or translate plan usage into API-token cost.

Plan access is also not a stable product identity. OpenAI’s Help Centre, updated 2 October 2026, says Codex is included across ChatGPT plans while Codex Cloud eligibility varies by plan, rollout and workspace settings; it also says usage varies by plan and task factors. The benchmark should record the actual account and workspace state visible during each run, not assume behaviour from a plan name or promise that another reader has the same access.

Triggers define provenance, not quality

For a manual cohort, the documented trigger is @codex review. For an automatic cohort, the repository must be connected and the relevant permission and settings must permit automatic reviews. The benchmark record should identify which route was used, whether a review appeared, and its timestamps. It should not turn this section into setup advice: connection procedures, interface controls and availability may change, and the purpose here is to preserve experimental provenance.

A review that does not appear is a negative operational result. Do not repeatedly retrigger it until one favourable output arrives and then discard the attempts. Instead, apply a pre-registered retry rule—for example, one retry only after confirming that GitHub and account state meet the study’s recorded prerequisites. Preserve both attempts and classify the original as a missing-review event. OpenAI’s status page is only a service snapshot; even if it reports normal operation, it is not a service-level agreement or proof that an individual run should have completed.

Use OpenAI’s internal result as a design clue, not a borrowed score

In an article published 20 July 2026, OpenAI described an internal evaluation containing known rule violations and safe counterexamples. It reported that a rule-guided variant recovered 98% of required custom findings in its named primary suite, compared with 58.3% for its baseline control. OpenAI assessed coverage, restraint, retention and actionability.

Those figures are not a prior score for this benchmark. They belong to OpenAI’s internal suite, rule variants and recorded conditions. They cannot be transferred to a reader’s repository, account, plan, unexposed model, future product version or chosen rules. They do, however, support a more disciplined measurement design than counting all comments as successes.

What the internal evaluation supports—and what it does not
Dimension What it supports for protocol design What it does not support Practical benchmark procedure
Coverage Measure whether required custom findings are recovered. No claim that this repository will approach 98%, nor that every class of defect is covered. Create pre-labelled consequential violations and calculate defect-level recall with numerator and denominator.
Restraint Include safe counterexamples so silence can be a correct result. No guarantee that scoped rules eliminate false positives. Measure comments on clean cases, false-positive comments per clean PR and human dismissal time.
Retention Check that custom guidance does not crowd out ordinary serious-defect findings. No evidence that generic P0/P1 recall remains unchanged in another corpus. Include a separate stratum of serious defects unrelated to the rules and compare recall by condition.
Actionability Judge whether findings contain enough evidence and a safe path for a reviewer. No guarantee that a correct comment is clear, correctly prioritised or safe to follow. Use blinded human scoring across location, consequence, evidence, safe path and severity fit.
Reported 98% versus 58.3% Shows why controlled variants can reveal a difference in a named suite. Does not set an expected lift, acceptance gate or general Codex benchmark. Choose local thresholds from repository risk and review capacity before any outputs are opened.

OpenAI’s evaluation guidance, checked 2 October 2026, recommends explicit objectives, representative datasets, metrics, logged results and automated scoring calibrated with human judgement; it warns against “vibe-based” evaluation. Its agent-evaluation guidance also distinguishes exploratory trace grading from repeatable datasets and evaluation runs. These are useful methodological distinctions, but they do not mean GitHub Code Review exposes the same trace tooling or that any proposed score has been obtained.

Because the Evals platform is retiring (see the evidence checkpoints), use durable repository and continuous-integration artefacts rather than making that retiring platform the mandatory runner. A versioned case manifest, raw GitHub exports, human labels and analysis scripts remain inspectable independently of a particular evaluation service.

Pre-register three conditions that answer different questions

Use the same labelled PR corpus across three conditions. Duplicate private benchmark branches or PRs so one condition’s comments do not contaminate another reviewer’s context. Randomise condition order, anonymise case identifiers for human scorers and preserve the exact base and head commit hashes. If product behaviour prevents equivalent repetitions, disclose that limitation rather than manufacturing symmetry.

  1. Condition A — no Code Review Rules: remove the proposed review-rule guidance while retaining any non-review fixture requirements that must remain constant. This estimates performance on the same cases without the custom rule treatment. It is not “no repository context” unless all other context is demonstrably identical.
  2. Condition B — deliberately broad root wording: place a broad version of the target policy in the root guidance. This is a noise stress condition, not a straw man to be made nonsensical. It tests whether underspecified repository-wide wording creates comments on safe paths or unrelated files.
  3. Condition C — three concise scoped rules with a safe path: express the same intended outcomes within the relevant path scope and state when the pattern is permitted or how to handle it safely. This is the candidate treatment. Rules must be versioned, and the applicable AGENTS.md blob hash must be recorded for every run.

For example, suppose a repository requires privileged configuration changes to pass through one validated helper, while test fixtures may construct the same data directly. Condition B might state the broad policy at the root without distinguishing the fixture path. Condition C would limit the rule to production configuration paths and identify the approved helper and safe test exception. The benchmark would then contain both genuine bypasses and valid fixture constructions. This is an example of experimental structure, not recommended rule wording or a prediction that C will perform better.

Pre-register hypotheses that can fail

Primary hypothesis: condition C increases defect-level recall for consequential repository-specific violations relative to condition A, while satisfying all guardrails. “Increases” must be represented by a numerical gate fixed in advance, not merely a positive point estimate.

Rule-design hypothesis: condition C produces a lower clean-PR false-positive burden than condition B. This tests the practical value of scope and safe-path wording; it does not assume B must fail.

Retention hypothesis: condition C is not worse than condition A beyond the pre-registered non-inferiority margin for ordinary P0/P1 defect recall. A custom-rule gain is not acceptable if ordinary serious defects disappear from review.

Human-effort hypothesis: condition C does not exceed the owner’s pre-agreed median triage-time increase and does not create an unacceptable upper tail. Median time alone can hide a small number of extremely costly reviews, so publish the distribution and case-level values.

Actionability hypothesis: the proportion of correct condition-C findings passing every required actionability component meets the pre-registered gate. Do not average a strong location score with a failed safe-path score if the rubric defines the latter as essential.

Separate the primary metric from guardrails

The primary metric should be consequential repository-rule defect recall: correctly recovered labelled rule defects divided by all eligible labelled rule defects. Count at defect level. If two comments identify one defect, the numerator increases by one, while the extra comment enters duplicate and comment-load metrics.

Guardrails should include ordinary P0/P1 retention recall, comment-level precision, false-positive comments per clean PR, proportion of clean PRs receiving any false-positive comment, duplicate rate, actionability pass rate, blinded human triage minutes per PR, time to first review, time to completed review, missing-review rate and visible usage proxy. Show every numerator and denominator. Where sample size permits, report an interval such as a bootstrap interval and retain case-level scatter rather than publishing only a rounded percentage.

Decision rule: condition C cannot pass by compensating one failed guardrail with exceptional primary recall. A rule set that catches more custom violations but loses ordinary serious defects, floods safe examples or consumes unacceptable reviewer time is a failed rollout candidate under this protocol.

Fix exclusions, failures and stop conditions before runs

An exclusion policy protects the dataset from genuine corruption; it must not become a mechanism for removing inconvenient outcomes. Define permitted exclusions narrowly. Examples include a PR whose head commit does not match the manifest, accidental exposure of a gold label to the reviewer, a corrupted diff, or a repository connection demonstrably absent before the trigger. Preserve excluded records, state who authorised exclusion and publish sensitivity results with and without disputed cases where safe.

Do not exclude a run because it produced no comments, took a long time, repeated a comment, disagreed with the gold label or failed an acceptance gate. A no-review event should remain in the operational completion denominator unless the pre-registration identifies a specific platform-invalidating condition. If OpenAI’s status page reports an incident, record the contemporaneous evidence; do not assume that the incident caused the failure or that an “operational” status disproves it.

A case failure occurs when required provenance is unavailable, the PR differs across conditions, the applicable rule revision cannot be established, or a secret or untrusted production datum has entered the fixture. A condition failure occurs when the run count cannot meet the pre-registered minimum, blinding is broken systematically, or the product surface changes in a way that prevents meaningful comparison. A pilot failure occurs when condition C misses a decision gate, causes a guardrail breach or cannot be operated within the safety boundary.

Stop immediately for exposed credentials, customer data, private incident information, unauthorised source code, unexpected write access, unrestricted outbound network access, attempted auto-merge, or any action outside the approved review-only boundary. Revoke affected credentials where applicable, follow the organisation’s incident process and do not publish contaminated artefacts. Keep untrusted data and all secrets out of prompts and rule files. Use a public fixture or synthetic sanitised mirror, least privilege and no production credentials.

OpenAI’s Codex security documentation distinguishes sandbox boundaries, approval policy and network controls, and says monitoring does not replace sandboxing, permissions or result review. These controls are benchmark inclusion requirements, not evidence that the resulting review is secure. Preserve tests, branch protections and required approvals. Never allow benchmark comments to merge code automatically.

Data controls must also be disclosed before collection. OpenAI states that content from individual ChatGPT and Codex services may be used to improve models unless the relevant user opts out, while inputs and outputs from ChatGPT Business, Enterprise, Edu and the API are not used for model improvement by default. It separately describes a Codex “Include environments” control. These statements do not resolve an organisation’s contractual, source-code, privacy or customer-data duties. The decision owner must confirm the applicable account type, workspace policy and data settings before any run.

A worked hypothetical registration record

Warning: the gates below are illustrative only. They are not product guarantees, industry standards or universal recommendations. A repository handling payments, identity, healthcare, employment or government decisions may require stricter gates, specialist review or exclusion from this pilot. Thresholds must reflect the repository’s severity policy, reviewer capacity and risk appetite and must be approved before results are viewed.

Example pre-registration with intentionally blank result fields
Decision owner Developer-experience lead; named approver recorded in the private registration
Decision outcome Whether to retain condition C for a larger advisory-only pilot in this repository
Surface Connected-GitHub Codex Code Review; manual @codex review cohort only
Fixture Sanitised repository; pinned base commit, lockfiles and test command; no production access or secrets
Conditions A: no Code Review Rules; B: broad root wording; C: three concise scoped rules with documented safe paths
Primary population Pre-labelled repository-specific violations meeting the written P0/P1 policy
Primary gate Condition C must improve absolute rule-defect recall over A by at least the locally approved margin: example margin, 15 percentage points
Retention gate Condition C ordinary P0/P1 recall must not fall below A by more than example margin, 5 percentage points
Noise gate Condition C must produce fewer than example threshold, 0.25 false-positive comments per clean PR, and must be lower than B
Time gate Condition C median blinded triage time must not exceed A by more than example threshold, two minutes per PR; upper-tail cases receive separate review
Actionability gate At least example threshold, 80% of correct C findings must pass all mandatory rubric components
Exclusions Commit mismatch, corrupted fixture, confirmed gold-label exposure or missing repository connection established before trigger; all exclusions retained in the audit log
Immediate stops Secret or private-data exposure, unexpected write capability, unrestricted egress, auto-merge attempt, broken blinding across the cohort or unrecorded rule change
Primary result _____ / _____; estimate _____; interval _____
Retention result _____ / _____; difference from A _____
Noise result _____ false-positive comments / _____ clean PRs; distribution _____
Triage-time result Median _____ minutes; range or interval _____; excluded timing records _____
Visible usage proxy Account/workspace _____; visible measure _____; no API-price conversion
Decision _____ retain for advisory pilot / rewrite and rerun / do not expand

Apply the gates conjunctively: all mandatory gates must pass. If the recall point estimate rises but its uncertainty is too wide for the owner’s decision, the honest outcome is “insufficient evidence”, not “promising”. If C beats A but not B on noise, the scoped wording has not demonstrated the intended restraint advantage. If C improves custom-rule recall while breaching ordinary-defect retention, reject or narrow the rules. If quality passes but triage burden exceeds capacity, keep the review advisory, reduce rule scope or decline rollout.

Do not repair a failed result by editing labels, dropping difficult safe counterexamples or selecting only the best repetition. A post-run rule rewrite creates a new benchmark version and requires a new pre-registered run on an untouched or separately identified holdout. Preserve raw negative results, reviewer disagreements and missing reviews. That audit trail is part of the evidence and must not be removed.

The output of this section is therefore a signed registration, not a conclusion about Codex. It names the owner, product boundary, conditions, hypotheses, primary metric, guardrails, exclusions, safety stops and decision gates before the benchmark can influence them. Only later execution against controlled PRs can fill the blank fields, and every security-, privacy-, money- or other consequential interpretation must remain subject to qualified human review.

Build a labelled pull-request corpus before connecting the reviewer

The examples below specify fixtures and labels that a team could implement; they are not product outputs or guarantees. Any eventual claim must be tied to the released repository, exact commits, rule blobs, run records and human judgements that produced it.

The benchmark repository should make one narrow comparison possible: does repository-specific AGENTS.md guidance improve detection of consequential, pre-labelled repository-rule violations without sacrificing detection of ordinary serious defects or imposing unacceptable false-positive and triage burdens? A collection of memorable pull requests is insufficient. The corpus needs matched counterexamples, defects unrelated to the rules and realistic distractions so that a reviewer cannot score well merely by repeating rule language whenever it sees a familiar file.

OpenAI’s custom-rule article, published on 20 July 2026, describes an internal evaluation built around known violations and safe counterexamples, with coverage, restraint, retention and actionability as distinct dimensions. The reusable methodological distinction is that detecting a seeded violation tests coverage, leaving a valid counterexample alone tests restraint, retaining ordinary defect detection tests retention, and explaining a usable remedy tests actionability.

Three parallel pathways pass the same sets of coloured blocks through differently shaped rule gates before reaching separate collection trays.
Holding the pull-request cases constant across no-rule, broad-rule, and scoped-rule conditions makes the proposed comparison inspectable.

Construct one disposable fixture, not a sanitised-looking production mirror

Create a dedicated GitHub repository whose entire purpose is evaluation. Prefer synthetic code with small, comprehensible services over a copy of an internal repository with names changed. Renaming identifiers does not necessarily remove customer data, proprietary logic, credentials, incident details or commercially sensitive architecture. If a mirror is unavoidable, require a documented redaction review by the repository owner, security representative, privacy owner and relevant administrator before connection.

The fixture must be runnable without production systems. Pin the base commit, dependency lockfiles, test command, operating-system or container-image identifier, compiler or runtime versions, and any GitHub Actions workflow needed to establish test status. Replace external dependencies with local fakes or inert test doubles. A test double is a controlled substitute for a real service; it should produce deterministic responses without network access or real accounts. Do not include production credentials, copied tokens, private keys, customer payloads, support transcripts or live endpoints, even if a test is expected not to use them.

Version every benchmark-affecting artefact in Git. That includes source files, fixtures, lockfiles, tests, case-generation scripts and each AGENTS.md variant. Do not rely on an untracked instruction pasted into an interface. For each rules condition, record the Git blob Secure Hash Algorithm (SHA)A family of cryptographic hash algorithms standardised for producing fixed-length message digests. Open glossary entry of every applicable root or nested AGENTS.md file. A blob SHA identifies the exact stored file contents; a commit SHA alone is not enough when investigators later need to confirm which scoped guidance covered a changed path.

OpenAI’s GitHub documentation, checked on 2 October 2026, says applicable root and more-specific nested AGENTS.md guidance can govern changed files. It also describes @codex review and automatic reviews for connected repositories, with comments focused on P0/P1 issues. These product facts determine what the fixture must record, but they do not establish that a particular rule will be followed or that a defect will be found.

The repository should contain a machine-readable build manifest. As an example, it might declare a pinned container digest, npm ci as the dependency installation command and npm test -- --runInBand as the test command. This is only a fixture design example, not a required Codex configuration. If a fresh authorised evaluator cannot recreate the base state and test result from the published materials without contacting a production service, the fixture is not ready.

Allocate 64 pull requests across four disclosed strata

Use 64 primary cases, divided evenly into four strata of 16. A stratum is a declared group of cases designed to test one aspect of behaviour. Balance prevents an easy class from dominating the headline result and makes the denominator obvious. It does not make 64 statistically universal, nor does it guarantee narrow uncertainty intervals. Report uncertainty and every case-level result rather than treating the round number as proof of representativeness.

  1. Sixteen seeded repository-rule violations. Each case should contain one primary, consequential breach of a proposed repository invariant. An invariant is a condition the repository expects valid changes to preserve. Examples include a required authorisation check before a privileged state transition or a mandated redaction path before audit data is emitted. Seed the smallest plausible defect that violates the rule, and avoid placing the rule’s own wording in the pull-request title or source comment.
  2. Sixteen valid safe counterexamples. Match each violation with a superficially similar change where the invariant is satisfied through an approved path. The pair should share vocabulary, paths and diff shape where practical. This tests whether guidance distinguishes semantics from keywords. A warning on the safe member is a false positive unless blinded reviewers identify a different valid defect.
  3. Sixteen ordinary high-severity defects unrelated to repository rules. These test retention: rules should not crowd out serious general review. The defects must not be restatements of the proposed AGENTS.md guidance. Suitable fixture categories may include a lost authorisation check, destructive operation against the wrong identifier, irreversible data truncation or a concurrency error that can corrupt durable state, provided each meets the pre-registered severity policy.
  4. Sixteen realistic mixed or noisy diffs. These should resemble review work rather than puzzle snippets. Combine benign refactoring, generated-file churn, tests and naming changes with either one labelled defect or, in designated clean cases, no primary defect. Record which form each case takes. Noise should challenge localisation and restraint without relying on hidden dependencies or deliberately malformed code.

Do not create all violations first and invent their safe counterparts later. Design each pair together. Start with an invariant, write the violating path and approved safe path, then make their surrounding changes comparable. Run the same tests against both. If the violation is supposed to pass existing tests because the repository lacks coverage for that invariant, state that explicitly; if it should fail a test, preserve that result. Test status is evidence about the fixture, not the gold label by itself.

A practical balance within the 16 noisy cases is to pre-register which are defect-bearing and which are clean. For example, a team could choose eight with one P0/P1 defect and eight with none. That allocation is a suggested design choice, not a product requirement. Fix it before reviews begin. Otherwise, evaluators may quietly add clean cases after seeing too many comments or remove difficult cases after observing misses.

Give every primary defect a written P0/P1 basis

Because OpenAI’s GitHub documentation describes Code Review comments as focused on P0/P1 issues, a benchmark that labels style concerns, minor maintainability problems or speculative improvements as required findings would misalign its primary truth with the documented surface. Every primary violation and ordinary defect therefore needs a written severity rationale established before any run.

Define the local policy rather than assuming that “P0” and “P1” have universal meanings. One defensible fixture policy could define P0 as a defect capable of causing immediate, broad and difficult-to-reverse compromise, data loss or service failure in the represented system. It could define P1 as a defect with a credible path to material unauthorised action, durable data corruption, major outage or another severe operational consequence that requires prompt correction. These are example benchmark definitions, not OpenAI guarantees and not a substitute for an organisation’s incident taxonomy.

The severity record should name the affected asset, triggering conditions, plausible consequence, expected scope and why existing controls do not make the path harmless. “Security issue” is not a basis. “A caller lacking the administrative capability can reach the destructive transition because the new handler bypasses the shared authorisation guard; the fixture’s route registration exposes that handler to authenticated non-admin users” is testable. Reviewers may later disagree with the label, but they need a concrete proposition to assess.

Exclude a candidate from primary recall if its seriousness depends on undocumented production topology, secret business context or a contrived input not represented in the fixture. It can be retained as an exploratory case, clearly separated from the primary corpus. The decision rule is: a competent reviewer should be able to verify both the defect and its severity basis from the sanitised repository, diff and declared assumptions.

Publish a case manifest that can reconstruct the comparison

Maintain the authoritative manifest in a stable format such as comma-separated values or JavaScript Object Notation (JSON)A text format for representing structured data as objects, arrays, numbers, strings, and other values. Open glossary entry, then render a readable table. One row represents one pull-request case, not one run or comment. Repeated runs belong in a separate run manifest keyed to case_id. The following rows are synthetic examples of schema use; they do not report observed Codex behaviour.

Example case-manifest schema and illustrative synthetic rows
Case identifier (ID)A value used to distinguish one record, task, source or object from another. Open glossary entry Base SHA Head SHA Changed paths Stratum Anonymised issue description Expected location Test result Provenance Label-release state
RV-07 <40-hex-base> <40-hex-head> services/ledger/refund.ts, tests/refund.test.ts Seeded rule violation Privileged refund transition bypasses the repository’s approved authorisation wrapper refund.ts, changed handler body Fixture tests pass; invariant-specific assertion intentionally absent Synthetic, authored from pre-registered invariant INV-AUTH-02 Encrypted gold; release after run lock
SC-07 <40-hex-base> <40-hex-head> services/ledger/refund.ts, tests/refund.test.ts Safe counterexample Same transition uses the approved wrapper and propagates its denial result No primary defect expected Fixture tests pass, including denial-path assertion Synthetic matched pair for RV-07 Encrypted gold; release after run lock
OD-11 <40-hex-base> <40-hex-head> workers/archive.ts Ordinary high-severity defect Deletion selects records by tenant display name rather than immutable tenant identifier archive.ts, changed deletion predicate Existing unit tests pass because fixtures use identical display-name and identifier values Synthetic defect independent of repository rules Encrypted gold; release after run lock
MX-04 <40-hex-base> <40-hex-head> workers/archive.ts, generated schema, tests and documentation Mixed/noisy diff Clean refactor with generated churn; no primary defect No primary defect expected All fixture tests pass Synthetic transformation plus regenerated artefacts Encrypted gold; release after run lock

In the real manifest, use actual immutable SHAs rather than placeholders. Store changed paths as a deterministic list, including renamed and deleted files. The expected location should be precise enough to judge localisation but robust to line-number drift: path plus function, symbol or changed hunk is preferable to a bare line number. Provenance should distinguish wholly synthetic creation, transformed open-source fixture material with permitted use, and any separately authorised natural-history case.

“Label-release state” should be operational, not decorative. Suggested values are draft, adjudicated, encrypted-gold, run-locked and public. Do not expose issue-bearing titles, branches or test names that reveal the hidden label to the reviewing condition. The public manifest can initially contain neutral descriptions or redacted fields, but the complete adjudicated record must be released after runs are locked if the team publishes benchmark claims.

Worked pair: a violation and a safe counterexample

Consider a synthetic fixture service that processes high-value refunds. Its repository invariant states that handlers changing refund state must call authoriseRefundTransition(actor, refund, targetState) and must return its denial without mutating state. The rule is scoped to services/ledger/. The fixture documents that bypassing this wrapper permits an authenticated but unauthorised operator to approve a refund, satisfying the pre-registered P1 basis.

In example case RV-07, the pull request adds an expedited endpoint. The changed handler loads the refund, checks only that the caller is authenticated, then calls refund.markApproved(). Tests cover the successful administrative path but omit a non-admin denial case, so the suite passes. Gold truth identifies one defect at the mutation hunk: the privileged transition bypasses the required wrapper. A qualifying finding would need to identify that bypass and its unauthorised-state-change consequence. Merely saying “consider adding validation” would not satisfy the actionability rubric.

In matched example SC-07, the diff is nearly identical in size and vocabulary, but the handler calls authoriseRefundTransition, immediately returns the denial result when authorisation fails, and mutates state only after approval. Its tests include both approved and denied callers. The gold label says no primary defect. A comment warning that the endpoint lacks authorisation because it sees markApproved() would be a false positive. A comment identifying a separate real P0/P1 defect could be credited only after blinded human validation as an incidental finding.

The pair tests a meaningful distinction: recognising whether the approved safety path is actually present, rather than reacting to a sensitive operation. It also prevents a misleading success criterion in which the reviewer “finds” the rule on every refund change. If both members attract the same warning, coverage may look high while restraint is poor.

Worked pair: an ordinary defect and a noisy clean diff

For an unrelated ordinary defect, use a synthetic archival worker outside the scope of the refund rule. In example OD-11, a refactor changes a deletion predicate from immutable tenantId to user-editable tenantDisplayName. Two tenants can share a display name, so one tenant’s archival job can delete another tenant’s durable records. Existing tests pass because their fixtures use the same string for both fields. The written P1 basis identifies cross-tenant destructive deletion, the collision condition and the absence of a later recovery step.

This case tests retention. Conditions with repository guidance should still detect serious defects that the guidance does not mention. If rule-guided review finds more seeded refund violations but misses this deletion flaw more often than the no-rules condition, the trade-off belongs in the result. Do not redefine the deletion bug as “out of scope” after observing it; ordinary P0/P1 retention is a pre-registered guardrail.

Match it with example MX-04, a clean but visually noisy archival refactor. Rename a local variable, extract a pure helper, regenerate a schema file and update tests, while preserving selection by immutable tenantId. Include enough generated churn to make the sensitive deletion call less visually dominant, but document the generator and reproduce its output. The gold record expects no primary defect. This case asks whether a reviewer can refrain from alleging cross-tenant deletion merely because it sees the same worker and deletion vocabulary.

These examples are proposed fixture content, not demonstrations of what Codex will say. Do not put the anonymised issue description in the pull-request title. Suitable neutral titles might be “Update archival worker” and “Adjust refund workflow”, randomly assigned from a pre-generated title list. Preserve the mapping privately until labels are released.

Make defect-level truth the primary unit

The unit of recall is the labelled defect, not the number of comments. A defect is counted as found once when at least one review comment identifies the affected behaviour, locates it adequately and explains the material consequence or approved safe path to the standard fixed in the human rubric. Three comments describing the same missing authorisation check do not become three true positives.

Duplicates still matter. Mark additional comments that substantially repeat an already credited finding as duplicates. Count the defect once for recall, but include every duplicate in comment volume, duplicate rate and reviewer triage time. This separates “did the review recover the defect?” from “how much material did a human have to process?” A summary and an inline comment may be duplicates if they assert the same defect without adding a distinct decision-relevant issue.

A comment on a safe counterexample or clean noisy case is a false positive when its claimed defect is not present. Weak phrasing alone does not automatically make a true observation false; it may instead fail the actionability threshold. Preserve separate labels for factual validity and actionability so that a correct but unusably vague comment does not disappear into either category.

Allow distinct valid incidental findings, but do not let them rewrite the primary gold set. At least two blinded reviewers should determine whether an incidental comment identifies a real, separate P0/P1 defect supported by the fixture. If accepted, record it as valid_incidental, give it a new defect identifier and report it separately. Do not retroactively add it to the primary recall denominator for selected conditions. Otherwise, a condition that emits many speculative comments could gain an unfair recall advantage through post hoc label expansion.

Use an unclear category when the fixture lacks enough evidence to decide. Resolve it through adjudication where possible; if unresolved, publish it with the exclusion rule chosen before analysis. Never silently count unclear comments as false positives in one condition and true findings in another.

Hide gold labels during runs, then release them

Keep the complete gold file outside the connected repository until collection ends. Encrypt it or place it in an access-controlled repository, record its SHA-256 digest and timestamp that digest before the first trigger. SHA-256 is a cryptographic hash used here to show that the label file was not changed after results became visible; it does not prove that the labels are correct. Limit access to the corpus designers and adjudicator, not the operators triggering reviews or the blinded comment assessors.

Gold labels include the defect identifier, stratum, severity basis, trigger conditions, expected location, acceptable semantic descriptions, safe path and pair identifier. Human assessors should receive de-identified comments and sufficient diff context, but not condition A, B or C, rule version, run number, issue-bearing branch name or the author’s expected verdict. Separate annotation packets can contain the adjudicated defect reference without revealing which experimental condition generated a comment.

After all runs, exclusions and annotations are locked, release the plaintext label file, its pre-run hash, adjudication changes and final hash. Explain every difference between the committed pre-run labels and the released form. Permissible changes might include correcting a path typo or documenting an adjudicated ambiguity; they must not include removing missed defects merely to improve a score.

Keep any natural-history holdout separate

A later holdout may use authorised, sanitised historical pull requests whose outcomes arose independently of the benchmark designers. Such cases can test whether synthetic fixtures omitted realistic interactions. They also introduce different selection, privacy and verification problems: historical review may reveal the original defect, repository evolution may break reproduction, and permission to publish code or comments may be absent.

Report that holdout separately by provenance and date range. Do not blend synthetic and natural-history cases into one headline recall, precision or false-positive figure. A combined number conceals their different sampling processes and can be moved by changing the mix. If permissions or redaction prevent release of case-level evidence, describe the holdout as non-public and avoid using it as the sole basis for a rollout claim.

Duplicate execution surfaces without changing the case

Create separate private duplicate pull requests or condition-specific branches from identical base and case head content. Conditions may differ only in the pre-registered rule treatment: no review rules, broad root guidance, or concise scoped guidance. Verify source-tree equivalence with commit-tree hashes where the rules arrangement permits it, and document unavoidable differences such as the presence of an AGENTS.md file.

Randomise condition order within each case so that A does not always run before B and C. Generate the order before collection with a recorded seed, then preserve the allocation file. Randomisation cannot remove product changes over a long collection period, but it reduces systematic confounding between condition and run order. Interleave cases rather than completing all no-rule runs days before all scoped-rule runs.

Run three repetitions per case and condition where the GitHub workflow permits independent repeated reviews without contaminating later outputs. A repetition must use a fresh duplicate pull request or another documented clean mechanism; asking again in a thread that already displays earlier review comments is not equivalent. If the workflow does not permit uncontaminated repetitions, record one run rather than simulating independence. Report the limitation.

Anonymise titles and branch names. Do not use missing-auth-check, safe-control or condition names visible to the reviewer. Use neutral, randomly assigned identifiers that do not expose stratum. Keep case IDs in the private mapping if even RV or SC prefixes would reveal truth.

Record enough provenance to explain every absence

For every attempt, record the Coordinated Universal Time (UTC)The internationally agreed standard for world time, used here to identify the time reference in timestamps and offsets. Open glossary entry start and end timestamps, case and repetition identifiers, randomised condition order, manual @codex review or automatic-review surface, repository connection state, base and head SHAs, applicable AGENTS.md blob hashes, changed paths and visible account or workspace metadata. UTC prevents local-time ambiguity. Capture visible platform or model metadata exactly as shown, but do not infer a model identity when the surface does not expose one.

Record visible settings through an export where available or a sanitised screenshot where permitted. Screenshots must not expose user email addresses, private repository names, tokens or customer data. Also record whether the pull request was draft or ready, its visible labels, author type, review permissions and any bot-visible description. Metadata can alter context even when source commits match.

Log repository connection state immediately before each run: connected or disconnected, automatic review enabled or disabled, and whether the operator could invoke the documented trigger. Access and workspace settings can change. OpenAI’s plan documentation, updated on 2 October 2026, says availability and usage vary by plan, task and workspace factors; therefore a future report must describe the observed configuration without promising that another account has the same access.

Capture failures as recorded outcomes, not administrative by-products. Suggested failure codes include trigger not accepted, permission failure, review never appeared within the pre-registered observation window, duplicate review contamination, connection changed, GitHub outage indication, test setup failure and operator error. Preserve the visible error text after redaction. Cite a timestamped status observation as incident context, not proof that it caused an individual run to succeed or fail.

“No findings” is a complete result. Save the pull-request state, trigger evidence, review state, timestamps and confirmation that no inline or summary finding appeared. Distinguish a completed review with no findings from a missing, failed or timed-out review. Treating both as an empty comment list would turn reliability failures into apparent restraint and could inflate precision.

Store raw artefacts where authorised: complete diff, review summary, inline comments, review state, timestamps, links, continuous integration (CI)A software-development practice that automatically integrates and tests changes in a shared repository. Open glossary entry logs and test output. Preserve original Markdown or structured exports alongside redacted public copies. Hash each artefact and list it in a run manifest. OpenAI’s evaluation guidance, checked on 2 October 2026, recommends explicit objectives, representative data, metrics and logged results rather than “vibe-based” assessment. Durable Git and CI artefacts are preferable here to prescribing the Evals platform, which is retiring (see the evidence checkpoint).

Apply safety and privacy gates before corpus admission

Use this inclusion checklist for every case and rerun it whenever the fixture changes:

  • Sanitised code: confirm that source, history, diffs, tests, screenshots and review exports contain no customer data, private incident details or restricted proprietary material. Reject a case that cannot be safely released or independently inspected.
  • Least privilege: grant the benchmark identity only the repository and review permissions required for the fixed workflow. Keep unrelated organisations and repositories inaccessible. Do not broaden permissions merely to make automation convenient.
  • No secrets: scan the complete Git history and generated artefacts for tokens, credentials, private keys and sensitive configuration. Use inert placeholders. Never place untrusted pull-request text together with secrets in a prompt or reachable environment.
  • No production access: use local fakes and disposable data. The fixture must not reach production databases, deployment systems, customer tenants or internal administrative services.
  • No unrestricted egress: allow only connections demonstrably required by the chosen review surface and repository workflow. OpenAI’s security documentation distinguishes sandbox boundaries, approvals and network policy; monitoring does not replace restrictive permissions or human review.
  • No auto-merge: reviews remain advisory. Preserve tests, branch protections and required approvals. OpenAI’s GitHub documentation explicitly says repository guidance does not replace those controls.
  • Data-control disclosure: document account type, workspace ownership, applicable data settings, retention decisions and who can access raw outputs. OpenAI’s data-use help page says individual ChatGPT and Codex content may be used to improve models unless the relevant control is used, while Business, Enterprise, Education and API inputs and outputs are not used for model improvement by default. That distinction does not remove contractual or organisational duties.
  • Legal and administrative review: obtain repository-owner, privacy, security, legal and workspace-administrator approval where applicable before connecting the service or publishing artefacts. Do not infer permission from technical access.

Human review remains mandatory for security, privacy, money, employment, government and other consequential decisions. Benchmark comments must not authorise a deployment, approve a financial transition, determine employment action, certify compliance or replace a qualified security assessment. The corpus may represent serious defects, but an automated finding or absence of one is not a security certification.

The final admission rule is strict: if a case cannot be reproduced without secrets or production access, cannot be severity-labelled from visible evidence, cannot be safely disclosed under the chosen data controls, or cannot be independently judged after labels are released, exclude it from the primary 64. Record the exclusion and reason. A smaller auditable corpus is more useful than a nominally complete one whose provenance or safety cannot withstand review.

Score signal, noise, human time and incomplete runs separately

Every threshold and number below is either a definition or an unmistakably synthetic teaching example. OpenAI’s documented GitHub review features establish what can be tested; they do not establish how the product will perform in a particular repository.

The scorecard should answer three different questions. First, did repository-specific guidance help the reviewer identify defects that the guidance was intended to cover? Second, did it preserve findings on serious ordinary defects unrelated to those rules? Third, what noise, delay, human effort and visible account usage accompanied the change? A single count of comments cannot answer any of these reliably: several comments may describe one defect, a clean pull request may attract plausible but incorrect warnings, and a run may fail to produce a review at all.

Calculate results for each pre-registered condition and compare paired observations from the same pull-request cases. Condition A has no proposed Code Review Rules; condition B has deliberately broad root guidance; condition C has concise, scoped guidance. Keep the trigger, repository connection, base and head commits, tests and account or workspace fixed as far as the product permits. A difference between conditions is interpretable only when the underlying case is the same and differences in execution provenance are disclosed.

Two people independently sort identical abstract tokens behind a divider, then bring disputed pieces to a shared central tray.
Blinded independent review and visible adjudication prevent a single evaluator’s confidence from becoming the benchmark’s ground truth.

Freeze the analysis record before opening the outputs

Create one analysis specification containing the metric name, unit, eligible cases, numerator, denominator, treatment of duplicates, treatment of unclear comments, missing-review policy, interval method and decision threshold. Commit that specification before reviewers see condition identities or aggregate scores. This prevents a team from redefining a miss, removing an inconvenient case or selecting a favourable average after learning which condition performed best.

Use stable identifiers at three levels. A case_id identifies the underlying pull-request diff. A run_id identifies one execution of one condition on that case. A comment_id identifies an individual review comment. Add a defect_id for each pre-labelled defect and for any distinct valid defect discovered during human review. These keys make it possible to distinguish a defect found three times from three separate defects.

The minimum long-form analysis table should have one row per run, with linked comment-level and defect-level tables. Record the condition, repetition number, start timestamp, first-review timestamp, completion timestamp, completion state, number of comments, matched labelled defects, independently validated additional defects, false positives, duplicates, unclear labels, triage minutes and visible usage change. Preserve “zero comments”, “no findings” and “no review appeared” as explicit states rather than blank cells.

Define a clean pull request before execution. For the primary clean set, use safe counterexamples and noisy diffs that contain no gold-labelled P0/P1 defect and are subsequently confirmed as clean under the blinded rubric. Do not quietly reclassify a case because the system made a convincing observation. If adjudicators validate a genuinely distinct serious defect, retain the original stratum in the manifest, mark the case as contaminated for the pre-registered clean-set analysis and report it separately. Dataset corrections may improve future benchmark versions, but they must not rewrite the current version’s primary result after outputs are known.

Defect recall measures coverage of known opportunities

Defect recall is the proportion of eligible gold defects that received at least one true-finding comment in a completed review. Its unit is a percentage or proportion:

defect recall = defects found at least once ÷ eligible gold defects

The denominator is defects, not pull requests and not comments. If one case contains two separately labelled defects, it contributes two opportunities. If three comments identify the same defect, that defect contributes one success. A comment counts as a match only when the blinded reviewers conclude that it identifies the affected location or code path and the substantive defect, rather than merely repeating a broad rule.

Report repository-rule recall on the seeded rule-violation stratum as the primary signal metric. Present separate per-rule counts so that a high-volume or easy rule cannot hide failure on another. For example, show R-01: found/eligible, R-02: found/eligible and R-03: found/eligible, as well as the pooled figure. If a defect is not covered by the applicable root or nested guidance for the changed path, it is not an eligible rule opportunity; classify that mismatch as a corpus or scoping defect rather than a review miss.

Comment precision measures whether emitted warnings deserve attention

Comment-level precision is the proportion of classifiable, substantive comments that are true findings:

comment-level precision = true-finding comments ÷ (true-finding comments + false-positive comments)

Report the unit as a percentage or proportion. The primary denominator excludes duplicates and comments left unresolved as unclear, but both excluded categories must be reported beside the metric. This prevents duplicate suppression from making the reviewer appear quieter than it was, and prevents uncertain comments from vanishing. Include a sensitivity analysis that treats every unclear comment as false, then another that treats every unclear comment as true. If the decision changes across those bounds, the evidence is not robust enough for rollout.

A true-finding comment can concern a pre-labelled defect or a distinct defect validated by both reviewers or adjudication. Keep those classes separate. Gold-defect precision answers whether comments matched known benchmark truth; expanded precision also credits legitimate additional findings. Do not add a post hoc defect merely because a comment sounds plausible. Require the same evidence and P0/P1 severity policy used for the original corpus.

Exclude non-substantive review boilerplate from precision only under a pre-written rule. Examples might include an empty completion notice or a summary that introduces inline comments without asserting another defect. If a summary makes an independently actionable claim, score that claim. Where one paragraph asserts two separable problems, split it into atomic claims before labelling and preserve the original comment identifier. The decision rule is whether a reviewer could accept or reject each claim independently.

Count false-positive burden where restraint matters most

False-positive comments per clean pull request measures warning burden on cases where the expected behaviour is restraint:

false-positive burden = false-positive comments on eligible clean PR runs ÷ completed eligible clean PR runs

The unit is comments per clean pull request. Report it separately for safe counterexamples and realistic noisy clean diffs. Safe counterexamples test whether a rule distinguishes a prohibited pattern from its permitted neighbour; noisy clean diffs test whether ordinary complexity provokes unrelated warnings. Pooling them alone could conceal a rule that behaves well on simple negatives but badly on realistic diffs.

Give the distribution, not just the mean. For each condition, publish the number and proportion of clean runs with zero, one, two and three-or-more false-positive comments, plus the median, interquartile range and maximum. A mean of one comment per clean PR can describe either uniform mild friction or a small set of severe cases with many false-positive comments. Those patterns require different action: rewrite a generally over-broad rule in the first case, or investigate path, language or diff-size interactions in the second.

Do not call a comment harmless merely because a human can dismiss it quickly. It still consumes triage attention and may affect trust in subsequent findings. Conversely, do not infer reviewer burden from count alone. A concise but subtle false positive may take longer to disprove than several obvious duplicates. Keep comment burden and timed human triage as distinct measures.

Protect ordinary-defect retention

Ordinary-defect retention recall is recall on pre-labelled serious defects unrelated to the repository-specific rules:

ordinary-defect retention recall = ordinary gold defects found ÷ eligible ordinary gold defects

Its unit is a percentage or proportion. “Ordinary” does not mean trivial: primary defects should satisfy the benchmark’s written P0/P1 policy because OpenAI’s GitHub documentation says review comments focus on P0/P1 issues. This metric asks whether adding special guidance coincides with lost attention to serious general defects. It does not measure all code quality or security coverage.

Compare condition C with A on exactly the same ordinary-defect cases. The pre-registered guardrail should state how much deterioration, if any, the team is willing to tolerate. A sensible method is to require no material reduction within a specified uncertainty bound, rather than declaring retention from equal rounded percentages. If the sample is too small to distinguish preservation from a consequential loss, report the result as inconclusive rather than “no regression”.

Condition B is diagnostically useful here. If broad rules recover custom violations but ordinary-defect retention falls, while scoped rules preserve it, scope may be the relevant intervention. If both B and C lose ordinary defects, added guidance may be competing with general review attention or the result may reflect run variability. The benchmark can reveal that pattern; it cannot establish an internal mechanism without product evidence.

Measure duplication and actionability rather than rewarding volume

Duplicate rate is the share of substantive comments that repeat a defect already represented in the same run:

duplicate rate = duplicate comments ÷ all substantive comments

The denominator includes true findings, false positives and duplicates, but excludes purely administrative boilerplate under the fixed rule. The unit is a percentage or proportion. Reviewers should mark a comment duplicate only when it describes substantially the same defect, consequence and affected path as an earlier comment. Similar rule citations attached to different defects are not duplicates.

Preserve ordering when identifying duplicates. The first adequate comment for a defect can be labelled true; later comments that add no materially different location, consequence or safe remedy are duplicates. If a later comment supplies necessary evidence that makes an earlier vague warning usable, the pair may be treated as one composite finding in a documented secondary analysis, but the primary comment-level table should retain both original labels.

Actionability pass rate is the proportion of true findings that satisfy every mandatory actionability field:

actionability pass rate = true findings passing all mandatory fields ÷ true findings assessed

The unit is a percentage or proportion. The mandatory fields are affected location, consequence, evidence, safe fix path and severity fit. Use an all-fields pass for the primary metric because averaging field scores can hide a critical omission: a comment with an exact line and detailed explanation is still unsafe to act on if its proposed path would violate the documented exception.

Also publish pass rates per field. That diagnostic table distinguishes comments that locate problems but omit consequences from comments that explain risks but propose no safe route. Treat a fix path as safe when it gives a bounded remediation direction consistent with the fixture’s constraints; it need not provide a complete patch. Reviewers must not execute suggested code merely to decide whether prose is actionable.

Measure latency and human work with different clocks

Time to first review output is elapsed wall-clock time from the recorded trigger event to the first visible substantive review output:

time to first output = first substantive-output timestamp − trigger timestamp

Use seconds or minutes, retain the raw timestamps in Coordinated Universal Time (UTC), and state whether the trigger was @codex review or automatic review. A GitHub event timestamp is preferable to a person’s stopwatch where available. If a review posts only a completion summary saying there are no findings, that completion is the first output. If no output arrives before the pre-registered timeout, record a right-censored or failed observation according to the fixed completion policy; do not enter zero.

Time to completed review is elapsed wall-clock time from trigger to the event that the protocol defines as completion:

time to completed review = completion timestamp − trigger timestamp

Define completion before runs. It might be the posting of a final GitHub review state after all inline comments, but the protocol must use only observable events and must not invent an interface state. Report median, interquartile range, minimum, maximum and case-level points for completed runs. Averages alone are vulnerable to long-tail delays, while medians alone hide operationally important outliers.

Do not describe lower system latency as human productivity savings. The tool can return quickly while producing comments that take longer to verify, or return slowly while requiring little triage. Queueing, service state, account limits and product changes can also affect latency. Record relevant status observations, but judge the run against the registered measurement window.

Blinded human triage minutes per PR is active reviewer time spent classifying and assessing the review output, excluding idle time and condition administration. Start timing when the reviewer opens the de-identified output; pause for interruptions; stop when all comment labels and actionability fields are complete. Its unit is active minutes per pull-request run.

Use the sum of both independent reviewers’ active minutes as the primary labelling-effort measure, because both reviewers contribute to the benchmark. Also report minutes by reviewer and the adjudicator’s additional time. Do not average away the second reviewer and then claim that the resulting number represents operational production review. The dual-review process estimates benchmark labelling effort; a future production workflow may differ and requires its own timed baseline.

Randomise packet order within practical constraints and prevent reviewers from seeing condition, rule text, bot identity cues or prior labels. Use the same viewing format for every condition. If redaction itself makes a comment harder to assess, flag that case and apply the declared exclusion rule. Record pauses and exclude unrelated interruptions, but do not remove legitimate investigation time such as inspecting the supplied diff or fixture tests.

Report completion and failure before quality scores

Every scheduled run must end in one of a finite set of states: completed with findings, completed with no findings, timed out with no visible review, failed with an observable error, cancelled by the operator, or protocol-invalid. Define these states in the registration. “No findings” is a valid completed output and may be correct on a clean case or a miss on a defect case. It must never be merged with “no review appeared”.

Publish scheduled, attempted, completed, failed, timed-out, cancelled and invalid counts by condition and case stratum. Give reasons for every exclusion. Operator mistakes, incorrect commit pairs, accidental disclosure of condition labels and repository-connection failures may justify invalidation if the same rule is applied without reference to output quality. Product failures remain outcomes for completion analysis even when they cannot be scored for comment quality.

The primary quality analysis should use completed runs because recall cannot be inferred from output that never arrived. Pair it with a conservative effectiveness analysis in which a missing review on a defect-bearing case counts as failure to surface the defect. On clean cases, a missing review must not be rewarded as perfect restraint. This two-view method distinguishes conditional comment quality from the end-to-end chance of obtaining useful review output.

For paired comparisons, a case contributes to the strict paired quality analysis only when all compared conditions have a scorable completed run for the same repetition block. Report the number of complete pairs and the unpaired missingness pattern. Also show all completed runs descriptively. If condition C has more missing reviews than A, complete-case precision can look attractive merely because difficult cases disappeared; the completion table prevents that distortion.

Respect pairing and repeated-run dependence

The same case reviewed under A, B and C is a matched set. Compare within-case differences before summarising across cases. For recall, calculate whether each eligible defect changed from missed to found or found to missed. For false-positive burden and triage time, subtract A from B and A from C for the same case and repetition block. This controls for fixed case difficulty more directly than comparing unrelated condition averages.

Three repeated runs of one case are not three independent pull requests. Product nondeterminism makes repeats useful for estimating stability, but multiplying the row count must not create false confidence. Treat case_id as the primary sampling cluster. Describe within-case variability, including cases that alternate between a finding and a miss, rather than reporting only pooled comment totals.

Use case-clustered bootstrap intervals as a practical uncertainty method. Sample case identifiers with replacement; for each selected case, carry all its conditions and repetitions into the resample; compute the paired metric difference; and repeat the process a pre-registered number of times. Report the central estimate and percentile interval, while disclosing the bootstrap procedure and random seed. The interval describes uncertainty for this constructed corpus under the analysis assumptions, not universal Codex performance.

For sparse binary outcomes, also publish raw numerator and denominator because intervals can be wide or unstable. Do not use “statistically significant” as a substitute for the operational gate. A small recall improvement may be too costly in false-positive comments, while an uncertain estimate may still expose a severe regression that warrants stopping. The pre-registered decision combines effect direction, uncertainty and practical thresholds.

Show case-level distributions through downloadable tables and plots generated from versioned scripts. Useful examples include paired dot plots for triage minutes, per-case false-positive counts, and a tile map showing found, missed, unstable or incomplete status for each defect across repetitions. These are suggested reporting methods, not promised product outputs. Avoid reducing the benchmark to one composite score: weights would obscure whether a gain came from recall, restraint or missing runs.

Worked synthetic calculation: arithmetic only, not a Codex result

The following invented miniature dataset exists solely to demonstrate denominators. It is not drawn from OpenAI, this publication, a customer repository or an executed pilot. Its tiny sample is unsuitable for a product conclusion. Every output cell is explicitly labelled accordingly.

Invented instructional inputs and derived values
Measure Condition A Condition C Arithmetic
Eligible repository-rule defects 10 — example arithmetic only 10 — example arithmetic only Fixed paired denominator — example arithmetic only
Repository-rule defects found 4 — example arithmetic only 7 — example arithmetic only A: 4 ÷ 10; C: 7 ÷ 10 — example arithmetic only
Defect recall 40% — example arithmetic only 70% — example arithmetic only +30 percentage points — example arithmetic only
True-finding comments 5 — example arithmetic only 9 — example arithmetic only Comment counts may exceed defects found — example arithmetic only
False-positive comments 3 — example arithmetic only 6 — example arithmetic only Human-adjudicated labels — example arithmetic only
Comment precision 62.5% — example arithmetic only 60% — example arithmetic only A: 5 ÷ (5 + 3); C: 9 ÷ (9 + 6) — example arithmetic only
False positives on four completed clean PRs 2 — example arithmetic only 5 — example arithmetic only A: 0.50; C: 1.25 comments per clean PR — example arithmetic only
Eligible ordinary defects found 7 of 8 — example arithmetic only 6 of 8 — example arithmetic only A: 87.5%; C: 75% retention recall — example arithmetic only
Duplicate comments 1 of 9 substantive comments — example arithmetic only 3 of 18 substantive comments — example arithmetic only A: 11.1%; C: 16.7% — example arithmetic only
Actionable true findings 3 of 5 — example arithmetic only 7 of 9 — example arithmetic only A: 60%; C: 77.8% — example arithmetic only
Median blinded triage time 4.0 minutes per PR — example arithmetic only 6.5 minutes per PR — example arithmetic only +2.5 minutes per PR — example arithmetic only
Completed scheduled runs 12 of 12 — example arithmetic only 10 of 12 — example arithmetic only Two C runs missing — example arithmetic only

The instructional lesson is not that C “wins”. In the invented arithmetic, C has higher rule-defect recall and actionability, but lower precision, more clean-case noise, lower ordinary-defect retention, longer triage and fewer completed runs. A pre-registered gate could reject it despite the recall increase. Conversely, a team must not reject it solely because its percentage-point recall gain differs from OpenAI’s internal suite; that external figure is not the benchmark target.

Missing C runs also make the completed-run comparison vulnerable to selection. The analyst should identify which case types were missing, calculate the conservative end-to-end defect result, and show complete paired cases separately. No confidence interval has been fabricated for this example because its invented observations are insufficiently detailed for a meaningful resampling exercise.

Apply a two-reviewer blinded rubric

Recruit at least two reviewers with relevant software-review experience. Disclose their expertise in aggregate: languages, framework familiarity, security-review experience where relevant, and familiarity with the repository pattern. Do not imply equivalence merely because both have senior titles. If one reviewer authored the fixture or labels, disclose that conflict and, where feasible, prevent that person from conducting the first independent output assessment.

Give each reviewer the same de-identified packet: the diff, necessary local context, applicable public fixture documentation, test output and review comments in a standard order or randomised order fixed by the protocol. Remove condition labels, rule-version clues and execution metadata that could reveal A, B or C. Do not remove technical context needed to judge correctness. Blinding is meant to reduce expectation bias, not make assessment artificially difficult.

For each atomic comment claim, reviewers complete these fields independently:

  • Affected location: pass if the comment identifies a sufficiently precise file, line, symbol or execution path for a developer to inspect the claimed problem.
  • Consequence: pass if it states a concrete failure, exposure or violated invariant rather than saying only that code is “unsafe”, “incorrect” or contrary to a rule.
  • Evidence: pass if the claim is supported by the diff and supplied context, including the relevant control flow or missing safeguard.
  • Safe fix path: pass if it offers a bounded remediation direction compatible with the fixture’s documented safe counterexamples and constraints.
  • Severity fit: pass if the consequence meets the written P0/P1 policy used to admit primary defects.
  • Finding class: exactly one of true finding, false positive, duplicate or unclear.

A true finding identifies a real, in-scope defect with adequate supporting evidence. A false positive alleges a defect that is absent, relies on contradicted assumptions or wrongly rejects an approved safe pattern. A duplicate substantially repeats another comment in the same review without adding a distinct affected location or consequence. Unclear is reserved for insufficient fixture context or genuinely ambiguous claims; it is not a compromise label for reviewer discomfort.

Require independent labels before discussion. Store each reviewer’s original record immutably, then compare. Disagreements go to a documented adjudication meeting or a third qualified adjudicator. Preserve the two initial labels, the final label, the reason for resolution and any corpus defect discovered. Never overwrite disagreement with consensus alone, because disagreement is evidence about rubric clarity and output ambiguity.

Report agreement for the categorical finding class and for each binary actionability field. At minimum publish raw percentage agreement and the confusion table. A chance-corrected coefficient may be added if its method and limitations are stated, especially where one label dominates, but it should not replace the raw counts. Report how many comments required adjudication and how much adjudication time was consumed.

Calibrate reviewers on a small practice set that is separate from the scored corpus. Discuss mismatches, revise definitions if necessary, freeze the rubric, and then begin independent scoring. If the rubric changes after scoring starts, version it and re-label all affected outputs without exposing condition identities. The decision rule is that identical claims must be judged under identical definitions.

Limit any large language model judge to calibrated assistance

A large language model (LLM)A machine-learning model trained on large text collections to understand and generate language. Open glossary entry judge may help structure comments, identify candidate duplicates or flag records for human attention only after calibration against a human-gold subset. OpenAI’s evaluation guidance recommends combining automated measurement with human judgement and calibrating automated scoring against human feedback. Model judges may also be sensitive to answer order or length. Test for those possible biases during calibration rather than attributing that particular warning to OpenAI’s evaluation guide.

Construct the human-gold subset before using the judge operationally. Include true findings, safe counterexamples, subtle false positives, duplicates, terse valid comments and verbose invalid comments. Randomise presentation order and create length-controlled variants where practical. Compare the judge’s labels with adjudicated human labels for each class and each actionability field. Publish the confusion matrix and disagreement examples; do not report only overall agreement.

If calibration is acceptable under a pre-set rule, use the judge only for a declared secondary analysis or triage aid. Humans must retain final authority over benchmark truth. Re-run a random human audit of judge-assisted records and all consequential disagreements. If changing comment order or verbosity materially changes labels, disable judge scoring for that field. Never send private source, secrets, credentials, customer data or untrusted executable content to an evaluation prompt.

Human review is mandatory for security, privacy, financial, employment, government and other consequential decisions. This benchmark does not certify code, establish legal compliance or justify automated merging. Treat model and human labels as evidence within the stated fixture, not as permission to deploy or make a consequential decision without the appropriate accountable reviewer.

Record usage as a capacity proxy, not an invented API bill

For every run, record the account or workspace used, its relevant settings, the trigger surface and any usage or credit indicator actually visible to the operator. Capture before-and-after values where the interface permits reliable attribution, along with timestamps and screenshots or exports that contain no private data. If no sufficiently granular indicator is exposed, record not observable; do not estimate tokens from diff size.

OpenAI’s Help Centre distinguished Chat, Work and Codex as of 2 October 2026. It stated that Codex use signed in through a ChatGPT account follows plan usage and billing, whereas use with an owner-supplied API key follows API pricing. Therefore a connected-GitHub Code Review pilot funded through a ChatGPT plan must not be converted into an API-token price per PR.

Call the measure visible plan-credit or usage proxy, not cost, unless the recorded account presents an attributable monetary charge for the exact run. Suitable units are the interface’s own displayed unit, a count of reviews against a documented capacity limit, or simply whether a visible allowance changed. State the account, workspace and observation date because plan access, limits, credits and rollout can change.

Do not call an included allowance free. It has opportunity cost and may constrain other work even when no incremental currency amount is visible. Equally, do not combine human minutes with a guessed salary rate to manufacture savings. A later financial analysis may use an organisation’s approved labour and capacity assumptions, but it must remain separate from this product-quality benchmark and receive human financial review.

Use interpretation patterns rather than a winner label

Pre-decision trade-off matrix
Observed pattern What it supports What it does not support Practical decision rule
High rule recall, high noise The guidance may surface more labelled rule defects. It does not show that wider rollout is worthwhile or that reviewers save time. Narrow or split the rules; proceed only if false-positive and triage gates remain within pre-set limits.
Low noise, lost ordinary defects The reviewer may be restrained on clean cases. Restraint does not compensate automatically for reduced P0/P1 retention. Reject or revise the treatment when the ordinary-defect guardrail is breached, even if precision rises.
Improved signal, increased triage More comments may be true and actionable. Higher quality does not establish productivity savings. Compare the added reviewer minutes with the team’s pre-registered capacity ceiling; consider advisory use or narrower scope.
Incomplete-run pattern The completed outputs may still be individually useful. Complete-case quality cannot establish dependable end-to-end performance. Pause rollout if completion falls below the fixed gate; investigate missingness without deleting failed runs.

A fifth pattern deserves explicit treatment: higher rule recall with unstable repetition. If the same defect is found in one run and missed in two, the pooled percentage may disguise unpredictability. Publish per-case hit sequences and set a stability rule, such as requiring a minimum proportion of repeated runs to find each critical defect. That threshold is an example method and must be selected before outcomes are viewed.

Another possible pattern is no material difference with wide intervals. The correct conclusion is “inconclusive on this corpus”, not “equivalent”. A team can enlarge the representative dataset, improve label clarity or stop because the evidence does not justify more reviewer attention. It must not lower the gate after seeing the results or cite OpenAI’s internal suite as proof that its own underpowered pilot would eventually succeed.

Where scoped condition C improves rule recall over A but performs no better than broad condition B on restraint or retention, the evidence does not support the specific scope design. Conversely, if B creates noise and C avoids it, report the paired differences and affected case types rather than claiming that concise rules are universally superior. The causal claim is limited to the recorded rule revisions, fixture and product conditions.

Keep governance controls outside the treatment

Tests, branch protections and required approvals must remain unchanged across A, B and C. OpenAI’s GitHub documentation says repository guidance does not replace those controls. They are safety boundaries, not variables to weaken in pursuit of a cleaner benchmark. Do not auto-merge, provide production credentials or allow unrestricted outbound network access. Keep untrusted data and secrets out of prompts and review packets.

If a condition causes a test or approval difference, classify the run as protocol-invalid unless that difference was part of the pre-registered design. Preserve the record and reason; rerun only under the documented retry policy. The benchmark concerns review-comment quality, not whether removing safeguards makes a workflow faster.

Finally, make the decision from the complete scorecard. Require the pre-set rule-recall improvement, ordinary-defect retention guardrail, clean-PR noise ceiling, actionability floor, triage-time ceiling and completion floor to be considered together. A condition that passes one headline measure while failing a consequential guardrail has not passed the pilot. Any rollout decision affecting security, privacy, money, employment, government services or other consequential activity requires accountable human review beyond this benchmark.

Publish an audit pack that another team can rerun

This protocol has not been run by this publication. It produces no measured recall, precision, time saving, cost saving, rating or recommendation for Codex Code Review. A credible publication begins only after the team can connect each reported number to a pull request, an immutable commit, a rule revision, a captured review and a disclosed scoring decision. A chart without that chain is a presentation, not an audit.

The durable unit of publication should be a versioned Git repository accompanied by local and CI scripts. The repository records what was tested; the scripts reconstruct the case inventory, validate run completeness and calculate the result table. This arrangement is preferable to making a hosted evaluation product mandatory because readers can inspect the inputs, preserve them after a service changes and rerun the analysis in their own controlled environment.

OpenAI’s evaluation-best-practices guidance, checked on 2 October 2026, recommends explicit objectives, representative datasets, defined metrics, logged results and automated scoring calibrated with human judgement rather than “vibe-based” evaluation. The Evals platform is retiring (see the evidence checkpoint). The protocol therefore retains those first-party evaluation principles without making the retiring platform a dependency. OpenAI’s agent-evaluation guidance, also checked on 2 October 2026, distinguishes exploratory trace grading from repeatable datasets and evaluation runs. That distinction supports reproducible benchmarking, but it does not establish that GitHub Code Review exposes the same traces or evaluation machinery.

Choose a repository layout that preserves evidence, not just conclusions

Create a separate audit repository or a clearly isolated directory in the sanitised fixture repository. Pin the audit release with a signed tag where organisational tooling permits. A practical example is:

README.md
environment/
  platform-record.yml
  settings-record.yml
  dependencies.lock
cases/
  cases.csv
  exclusions.csv
labels/
  labels-delayed.csv
  label-schema.json
rules/
  condition-a/
  condition-b/
  condition-c/
  revisions.csv
runs/
  manifests/
  raw-private/
  redacted-public/
diffs/
ci-logs/
rubrics/
  human-rubric.md
  calibration-cases.csv
  adjudication.csv
analysis/
  validate_pack.py
  build_results.py
  bootstrap_intervals.py
results/
  case-results.csv
  aggregate-results.csv
  publication-table.json
checksums/
  SHA256SUMS
decision/
  worksheet.yml
  signed-decision.md

This is an example structure, not a Codex requirement. Use another layout if it preserves the same relationships. A competent reviewer who was not involved in the pilot must be able to identify the exact input, condition, output, label and calculation behind every published cell. If a reported value depends on an undocumented spreadsheet edit, manual copy-and-paste operation or deleted GitHub view, the pack is incomplete.

Write the README as a reconstruction guide

The README.md should state the benchmark question, product boundary, fixture licence, corpus design, conditions, primary metric, guardrails, pre-registered thresholds, trigger method, repetition policy and publication date. It must say explicitly that the evaluated surface was connected-GitHub Codex Code Review rather than Codex Cloud, Codex Security, ChatGPT Work, ordinary ChatGPT or an OpenAI API workflow.

Document the commands needed to validate and analyse the pack, including interpreter and package versions. An example sequence might be:

python -m venv .venv
. .venv/bin/activate
pip install --require-hashes -r environment/dependencies.lock
python analysis/validate_pack.py
python analysis/build_results.py
python analysis/bootstrap_intervals.py

These commands are illustrative. Do not claim one-command reproduction if reviewers must obtain private data, restore deleted reviews or infer settings. If some raw artefacts cannot be published, describe exactly which calculations can be reproduced from the redacted release and which require controlled access. The decision rule is to label a result “independently reproducible” only when the released inputs suffice to regenerate it; otherwise call it “internally reproducible from access-controlled evidence”.

The README should also record the intended meaning of an absence. “No review appeared”, “review completed with no findings”, “artefact export failed” and “run was stopped by the safety policy” are different states. Collapsing them into an empty comment list would make a service or collection failure look like reviewer restraint.

Record the environment and settings without inferring hidden product details

The environment record should identify the fixture repository commit, operating system or container digest used for tests, dependency lockfile hash, relevant GitHub workflow revisions and test commands. The settings record should identify the account or workspace category, repository connection state, manual @codex review or automatic-review trigger, visible review settings, branch-protection state and any visible usage or product metadata.

OpenAI documents manual @codex review, automatic review for connected repositories, comments focused on P0/P1 issues and the use of applicable root and nested AGENTS.md guidance. As checked on 2 October 2026, those facts define the surface but do not establish its performance. Record what the interface exposed; do not infer a model identity, backend revision or reasoning configuration that was not displayed.

Capture settings as structured text where possible and screenshots only where permitted. Before publication, inspect screenshots for repository names, account identifiers, email addresses, tokens, customer data and unrelated browser content. Keep untrusted data and secrets out of prompts and fixtures. Never place production credentials, private incident details or unrestricted outbound access in the benchmark merely to make the environment resemble production.

OpenAI’s Help Centre, updated on 2 October 2026, says Codex access, Cloud eligibility and usage can vary by plan, rollout and workspace settings. It also distinguishes ChatGPT-account Codex usage from use of an API key under API pricing. Consequently, record the actual account path but do not convert visible ChatGPT-plan usage into an API-token bill. A future reader may be able to reproduce the method without having identical entitlement or capacity.

Make cases.csv the join key for the whole pack

Each row in cases.csv should have a stable case_id, stratum, base and head commit identifiers and blob identifiers (SHA values) as used by Git, changed paths, fixture version, intended trigger, test result, duplicate pull-request identifiers and a pointer to the delayed label. Include the disclosed severity policy and whether the case is eligible for the primary P0/P1 analysis.

Do not put an answer-revealing case title into a file available to the reviewer during execution. For example, RULE-VIOLATION-AUTH-BYPASS-07 leaks the intended defect. A neutral execution identifier such as case-027 can map to a descriptive label only after the runs close. Preserve that mapping in the delayed-label release.

Add a schema and validate uniqueness. Every eligible case must map to the planned conditions and repetitions; every run must map back to one case. The validator should fail, rather than silently drop a row, when a commit is missing, a path does not exist or a run references an unknown case. A warning is sufficient for an optional screenshot; a missing diff, condition or completion state should be a hard failure.

Release delayed labels without pretending they were always public

The labels package should contain the pre-run gold labels released after collection, plus evidence that those labels were frozen before outputs were inspected. Suitable evidence includes an encrypted file committed before the run with the decryption key released later, or a timestamped checksum recorded in an access-controlled register. The purpose is to make post-outcome relabelling detectable.

For each labelled defect, record its location, expected consequence, severity basis, acceptable finding boundaries and safe-path rationale. For clean counterexamples, state why the superficially similar code is valid. For ordinary defects, state why the defect is independent of the repository-specific rule treatment. If adjudicators discover a genuine defect absent from the original labels, preserve both records: classify it under the pre-registered policy for the primary analysis and, where justified, add a clearly marked post hoc analysis. Do not rewrite the original label history.

The decision rule is that primary results use only labels and eligibility rules fixed before outcome inspection. Newly discovered defects may inform a secondary analysis and the next corpus version, but cannot quietly enlarge one condition’s denominator or improve its apparent precision.

Track rule revisions as experimental treatments

The rules/ directory should contain the exact root and nested AGENTS.md files used for every condition, not paraphrases copied into the report. Record file path, Git blob hash, effective scope, revision identifier, parent revision, author, review approval and rationale in revisions.csv.

A change in wording, location or scope creates a new treatment version. For example, moving a rule from the repository root into a service directory can change which modified files it covers; it is not merely editorial housekeeping. Likewise, adding a safe-path exception after observing false positives is a protocol change. It may be sensible operationally, but the corrected condition must be rerun and reported as a new benchmark version.

Do not overwrite a failing rule revision and rerun only the cases it missed. Either retain the original result or begin the registered rerun policy for the new revision. The decision rule is that any treatment change made after outputs are visible ends comparability with the original pre-registered condition unless all affected cases are rerun under the new version.

Give every execution an immutable run manifest

One manifest per attempted run should contain the run identifier, case identifier, condition, repetition number, randomised order position, UTC start and end times, trigger type, pull-request identifier, base and head commits, rule blob hashes, visible product metadata, connection state, test status, review appearance status, collection status and exclusion status. Include pointers to the diff, raw artefact, redacted artefact and logs.

Use explicit enumerations rather than free text. A suggested example is:

{
  "run_id": "v1-case-027-c-r2",
  "case_id": "case-027",
  "condition": "C",
  "repetition": 2,
  "trigger": "manual",
  "review_state": "completed_no_findings",
  "collection_state": "complete",
  "eligible_primary": true
}

This sample illustrates structure only; it is not an observed run. Other review states might include completed_with_findings, no_review_observed, cancelled_by_operator and artifact_unavailable. The manifest must not convert all four into zero comments.

Validate run counts against the pre-registration. If a condition lacks planned repetitions, publish the gap before comparing quality. Rerun only under the declared retry policy. Otherwise, selective retries can favour a condition that happened to fail technically or produced an inconvenient result.

Preserve raw and redacted review artefacts

The private evidence store should retain the complete review summary, inline comments, timestamps, review state and stable links or export identifiers, where collection and retention are permitted. Preserve original Markdown or structured exports rather than only screenshots. Screenshots are useful corroboration but poor analysis inputs: they can crop context, obscure edits and resist automated validation.

Create a separately generated public-redaction layer. Redaction should remove secrets, personal data, private repository names and restricted source while preserving the text needed to assess location, consequence, evidence, severity and suggested safe path. Keep a redaction log that identifies the field category removed without reproducing the sensitive value.

OpenAI’s data-use help page distinguishes data controls across individual services and business, enterprise, education and API offerings. Because settings and policies can differ, disclose the account category and applicable controls used during collection, then obtain the organisation’s privacy and data-governance approval. Do not infer that a vendor default resolves contractual, customer-data, source-code or retention obligations.

If safe redaction would destroy the evidence needed to judge a finding, do not publish a misleading fragment. Mark the artefact access-controlled and explain how an authorised auditor can inspect it. Security, privacy, financial, employment, government and other consequential decisions require qualified human review; no automated score or redacted comment should be the sole decision-maker.

Archive diffs, CI logs and negative outcomes

Store the exact diff reviewed for each case, even when its commits are also in the fixture repository. This guards against later force-pushes, branch deletion or ambiguity about merge-base calculation. Record line endings and generated-file treatment if they affect matching.

Continuous-integration and test logs establish what the fixture’s mechanical checks did, not whether Codex found the defect. A passing test can coexist with a seeded defect if the corpus intentionally tests an uncovered invariant. A failing test can reveal that a case was malformed. Do not count a Codex comment as more correct because CI also failed unless the human rubric confirms it identified the labelled defect and consequence.

Publish “no findings” and clean runs with the same care as runs that produced comments. Negative artefacts are essential for recall and restraint. If the platform emits no review object, publish the manifest and collection evidence rather than manufacturing an empty review file that suggests successful completion.

Publish the rubric and adjudication history

The human rubric should define true finding, false positive, duplicate, unclear and out-of-scope comment, along with each actionability component: affected location, real consequence, adequate evidence, safe fix path and severity fit. Include reviewer instructions, calibration examples and the treatment-blinding procedure.

The adjudication file should retain each reviewer’s independent label, confidence where collected, disagreement category, adjudicated result, adjudicator and rationale. Never replace the independent labels with only the consensus result. Agreement is evidence about rubric stability; disagreement is evidence about ambiguity.

A meaningful trade-off appears when strict scoring excludes vague but directionally useful comments. Report that effect rather than loosening the rubric after inspection. For example, a comment that names a risky pattern but neither locates the defect nor explains the repository-specific consequence may be useful as a prompt for investigation, yet fail the registered actionability gate. The permitted conclusion is that it supplied advisory signal, not that it satisfied the benchmark definition.

Make analysis scripts simple, inspectable and deterministic

The analysis directory should contain scripts for schema validation, eligibility filtering, defect-to-comment matching, duplicate handling, metric calculation, interval estimation and table generation. Pin dependencies and expose random seeds used for bootstrap intervals. Prefer straightforward transformations over an opaque notebook with hidden state.

Generate publication tables from source files rather than manually entering figures. Include tests for denominator handling. For instance, comment precision is undefined when a condition emits no comments; it is not automatically 100%. Defect recall requires the number of eligible labelled defects, while false-positive comments per clean pull request requires the number of eligible clean pull requests. A script should refuse to merge those denominators.

Keep reviewer time, platform latency and visible usage proxies in separate fields. Human triage minutes are not wall-clock review latency; neither is a financial price. If visible usage data are unavailable or cannot be assigned to individual pull requests, report that limitation instead of estimating a per-PR cost.

List exclusions before presenting favourable subsets

The exclusions table should contain every planned and attempted case excluded from an analysis, the pre-registered reason code, the decision time, the person approving it and the analyses affected. Distinguish corpus invalidity, safety stop, product non-completion, artefact loss and protocol violation.

Never remove a technically incomplete run from the completion rate. It may be excluded from a quality denominator if the protocol says no review was available to score, but it still belongs in the operational record. Similarly, a malformed seeded case can be excluded from defect recall when adjudication shows the labelled defect was absent, yet its exclusion must be visible and applied consistently across conditions.

The decision rule is to publish both an intention-to-benchmark view, containing all scheduled runs and their completion states, and a quality-eligible view using the frozen exclusion policy. This separates operational reliability from comment quality without hiding either.

Checksum the release and generate a machine-readable result table

Generate SHA-256 checksums for every released artefact after redaction. Sign the checksum file where the team has an established signing process. Checksums show whether files changed after release; they do not prove that the original evidence was truthful or safely collected.

The machine-readable result table should have one row per condition, stratum and analysis version, with numerator, denominator, point estimate, uncertainty interval, missing-run count and exclusion count. Preserve case-level results in a separate table so independent analysts can inspect distributions and paired differences.

A useful schema includes:

benchmark_version
condition_id
rule_revision
stratum
metric_id
numerator
denominator
estimate
interval_method
interval_low
interval_high
eligible_cases
missing_runs
excluded_runs
analysis_commit
generated_at_utc

Do not put prose such as “excellent” or “acceptable” into measurement fields. Threshold comparisons belong in separate columns, such as gate_status and gate_definition_version. This allows readers to distinguish an observed estimate from the organisation’s policy judgement.

Publish results in layers rather than as one verdict

A responsible report separates measurements from interpretation. The following template is an example, not a product-generated report or a guarantee of available metadata.

1. Registered question and scope

  • State the repository, benchmark version, fixture commit and study dates.
  • Name connected-GitHub Codex Code Review as the surface.
  • Identify the trigger mode and the exact A/B/C rule revisions.
  • State the primary question and pre-registered thresholds.
  • List excluded surfaces, including Codex Cloud, Codex Security, ChatGPT Work and API workflows.

2. Observed measurements

  • Report scheduled, attempted, completed, scorable and excluded run counts.
  • Give numerators and denominators for defect recall, comment precision, false-positive burden, ordinary-defect retention, duplicate rate and actionability.
  • Report reviewer triage-time and review-latency distributions separately.
  • Report visible usage only in the units actually exposed.
  • Link each aggregate row to case-level and run-level evidence.

Use “observed in benchmark version X under recorded conditions”, not “Codex achieves”. The former describes bounded evidence; the latter incorrectly generalises across repositories, settings, future product versions and unavailable model details.

3. Uncertainty

  • Name the interval or resampling method.
  • Explain treatment of repeated runs and paired cases.
  • Show distributions and case-level variation, not only averages.
  • Report reviewer disagreement and adjudication volume.
  • Identify estimates too unstable for a rollout decision.

Do not use overlapping intervals as an automatic proof of “no difference”, or non-overlap as sufficient operational justification. Statistical uncertainty and policy thresholds answer different questions. A measured lift can still be operationally unacceptable if it creates too much triage work; a small uncertain lift can still motivate a larger pilot rather than deployment.

4. Limitations

  • State that seeded and synthetic cases may not represent natural pull-request prevalence.
  • Identify paths, languages, services and defect classes absent from the fixture.
  • Disclose incomplete runs, collection failures and redactions.
  • State that the benchmark addresses P0/P1-focused review under its severity policy, not comprehensive defect or vulnerability detection.
  • Explain that product nondeterminism and updates limit temporal generalisation.

5. Surface metadata

  • Record account or workspace category without asserting universal plan availability.
  • Record trigger, repository connection and visible settings.
  • Record only model or platform metadata actually exposed.
  • Describe branch protections, tests and approvals retained during the pilot.
  • State the data-control and privacy approval record.

6. Product or protocol changes

List observed interface changes, setting changes, documented product updates and treatment revisions by date. Preserve a timestamped status observation if it helps explain an incident; do not treat it as proof of an individual run’s health. If an incident plausibly affected runs, preserve those runs, label the context and apply only the registered retry policy.

7. Supported and unsupported interpretations

A supported interpretation must follow from the registered metric and scope. For example: “Condition C crossed the team’s pre-registered false-positive gate in this fixture version” is structurally supportable if the pack contains the calculation. “Codex is safe for autonomous merging” is unsupported regardless of the score because the benchmark did not test or authorise autonomous merging.

Other unsupported interpretations include security certification, universal repository performance, guaranteed time saving, a fixed price per review, superiority over untested tools, model-level claims when no model was exposed and reliability claims derived from a status-page snapshot.

Interpret patterns by changing policy, not by rescuing a preferred result

Higher rule recall with excessive clean-PR noise

This pattern means the rules may expose relevant violations while failing the restraint requirement. Inspect false positives by rule revision, changed path and safe-counterexample family. Determine whether one broad clause accounts for the burden.

The practical response is to narrow scope, add an explicit safe path or remove the noisy rule, then register a new revision and rerun the full affected corpus. Do not remove inconvenient clean cases. Until the revised condition passes the frozen gate, retain advisory-only use or decline rollout.

The permitted conclusion is bounded: the tested rule wording increased or preserved the measured finding signal but did not meet the organisation’s noise tolerance. It is not permissible to claim overall improvement merely because recall rose.

Low noise with weak repository-specific recall

A quiet reviewer can appear precise while missing the reason the rules were introduced. Check whether the rule was in applicable root or nested scope, whether the labelled defects met the severity policy and whether run manifests show completed reviews. This diagnosis distinguishes treatment failure from missing execution.

If execution was complete and scope was correct, narrow claims rather than celebrating restraint. The team may rewrite the rule, remove it because it adds no demonstrated value, or gather a larger representative corpus. The permitted conclusion may simply be “not enough evidence yet”, especially when intervals are wide or few eligible defects exist.

Repository-rule gains accompanied by ordinary-defect regression

This is a retention failure: attention to custom guidance may coincide with fewer recognised ordinary P0/P1 defects. First verify paired cases, completion and matching; then inspect de-identified comments for displacement, duplication or severity changes. Do not offset lost ordinary-defect recall by counting extra comments on rule violations.

The decision rule should follow the registered non-inferiority guardrail. If that guardrail fails, do not widen rollout even when the primary custom-rule metric improves. Suitable responses are to reduce the number or breadth of rules, keep the reviewer advisory, or stop the treatment pending a new protocol version.

Good comment quality with unacceptable reviewer effort

Actionable findings can still cost too much to triage. Examine the distribution of human minutes rather than only the median: a few complex or duplicated reviews may create queueing risk. Separate time spent assessing valid findings from time spent dismissing noise and resolving ambiguity.

If quality gates pass but the registered triage-time gate fails, the result supports neither a time-saving claim nor an unattended rollout. A team could limit use to selected repositories, paths or pull-request classes and measure that narrower policy as a new version. It could also retain manual triggering so a human chooses where the expected value justifies attention.

Quality appears acceptable but run completion is poor

Do not calculate quality only from successful runs and present it as the whole product experience. Publish completion separately and inspect trigger, connection, collection and contemporaneous status information. A small set of excellent completed reviews does not answer whether the workflow is dependable enough for the intended process.

If missingness differs by condition or case complexity, quality estimates may also be biased. The appropriate conclusion can be “not enough evidence yet”. Repeat only according to the pre-registered policy, or start a new benchmark version after correcting the protocol.

No threshold is crossed and uncertainty remains wide

An inconclusive result is a valid result. Check whether the corpus supplied enough eligible opportunities, whether reviewer disagreement was excessive and whether the thresholds demand more precision than the pilot could provide. Do not relabel cases, choose a more flattering metric or promote a secondary subgroup to headline status after seeing the data.

The next step may be a larger fixture release, clearer rubric or narrower decision question. Record that as a new registration. “Not enough evidence yet” is preferable to a rollout decision built from unstable estimates.

All registered gates pass

Passing gates supports only the rollout choice defined in the registration and only within the tested scope. It does not establish security certification, comprehensive bug detection or permanent performance. Begin with the smallest operational expansion consistent with the evidence, name an owner and set a reevaluation trigger.

Tests, branch protections, required approvals and qualified human review must remain in place. OpenAI’s GitHub documentation explicitly presents repository guidance as guidance rather than a replacement for those controls. Do not enable auto-merge as a benchmark conclusion. Security, privacy, money, employment, government and other consequential decisions require human review by appropriately authorised people.

Complete the reader worksheet before approving rollout

Copy this worksheet into the audit repository and complete it before opening condition labels or aggregate results. Values shown in square brackets are permitted options, not suggested answers.

benchmark_id:
benchmark_version:
decision_date:

scope:
  repository_or_fixture:
  fixture_commit:
  included_languages:
  included_paths:
  included_defect_classes:
  excluded_surfaces:
  github_trigger:
  rule_revisions:

registered_thresholds:
  primary_rule_recall_gate:
  ordinary_defect_retention_gate:
  comment_precision_gate:
  false_positive_clean_pr_gate:
  duplicate_rate_gate:
  actionability_gate:
  triage_time_gate:
  completion_gate:
  uncertainty_requirement:

privacy_and_safety:
  sanitised_fixture_confirmed:
  secrets_scan_completed:
  production_credentials_absent:
  outbound_access_policy:
  account_or_workspace_category:
  data_controls_recorded:
  retention_period:
  public_redaction_reviewed:
  privacy_approver:
  security_approver:

run_completeness:
  scheduled_runs:
  attempted_runs:
  completed_runs:
  scorable_runs:
  missing_runs:
  excluded_runs:
  retry_policy_followed:
  negative_results_published:

evidence_quality:
  labels_frozen_before_runs:
  conditions_blinded:
  two_independent_reviewers:
  disagreements_preserved:
  adjudication_complete:
  raw_artifacts_retained:
  redacted_artifacts_published:
  checksums_verified:
  analysis_reproduced:
  unsupported_model_inference_absent:

rollout_choice:
  decision: [remove_rules | revise_and_rerun | advisory_only |
             limited_rollout | not_enough_evidence]
  permitted_scope:
  controls_retained:
  owner:
  reevaluation_date:
  reevaluation_trigger:

Do not mark privacy or security approval complete merely because the fixture is synthetic. Review artefacts can still contain account details, contributor identities or model-generated reproduction of prompt content. Approval should cover collection, access, publication, retention and deletion.

For run completeness, reconcile totals mechanically. Scheduled runs should equal completed, missing, stopped and otherwise classified attempts under the registration. If the arithmetic does not reconcile, delay publication. For evidence quality, require a second person to execute the documented analysis from a clean checkout and compare generated checksums or tables.

Sign a decision record that distinguishes evidence from authority

The final record should be signed or formally approved according to the organisation’s normal engineering-governance process. A practical example follows:

Decision: advisory-only use for the stated scope
Benchmark version:
Evidence release commit:
Result table checksum:
Threshold document revision:

Observed gates passed:
Observed gates failed:
Indeterminate gates:
Material exclusions:
Material reviewer disagreements:
Known product or protocol changes:

Controls that remain mandatory:
- existing automated tests
- branch protections
- required approvals
- qualified human review
- no automatic merge
- existing incident and rollback procedures

Unsupported conclusions explicitly rejected:
- security certification
- comprehensive defect detection
- guaranteed reviewer-time reduction
- universal performance across repositories
- fixed API-equivalent cost
- inferred model performance

Decision owner:
Engineering approver:
Privacy approver:
Security approver:
Date:
Reevaluation trigger:
Signatures or approval references:

The rollout choice must correspond to the evidence. Choose remove rules when the tested guidance adds no defensible signal or creates unacceptable burden. Choose revise and rerun when a specific, remediable wording or scope problem is identified; the rerun becomes a new version. Choose advisory only when comments may assist humans but evidence does not support stronger process reliance. Choose limited rollout only when all applicable pre-registered gates pass for a clearly bounded scope. Choose not enough evidence when completion, sample size, disagreement, uncertainty or artefact quality prevents a decision.

Define reevaluation triggers concretely: a material rule revision, changed repository architecture, altered trigger mode, product-setting change, newly exposed product metadata, changed severity policy, sustained completion problems or a fixed calendar date. A status-page incident alone need not invalidate a release, but it can trigger review when it overlaps collection and plausibly affects completeness.

Conclusion

A trustworthy Codex Code Review pilot is not a gallery of plausible comments. It is a versioned account of what was eligible, what ran, what did not run, which rules applied, what humans judged and how each number was calculated. Repository and CI-based artefacts make that account inspectable after hosted tooling changes, while preserving OpenAI’s useful evaluation principles: explicit objectives, representative cases, logged execution, repeatable analysis and calibrated human judgement.

The resulting decision should remain narrow. Noisy rules should be narrowed or removed; useful but insufficiently proven output should remain advisory; protocol or treatment changes require a new version; weak evidence should produce “not enough evidence yet”. Even a passing pilot does not replace tests, branch protections, approvals or human review, and it is not a security certification. Keep secrets and untrusted data out of prompts, retain least privilege and require authorised human review for every consequential decision.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this