Build an Eval-Gated GPT-5.5 Batch Classifier: Strict JSON Outputs, Holdout Tests, Error Reconciliation, and Human Review


Evidence checkpoints
Documented point: A generative pre-trained transformer (GPT)A family of transformer-based language models trained to generate and analyze content. Open glossary entry-5.5 application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry batch classifier can use /v1/responses and Structured Outputs because the current model reference lists Batch and structured_outputs support. When a suitable snapshot is available, pin and record it, then retest after changes. [official source 1]
Documented point: The Batch API fits non-urgent full-export processing: JSON Lines (JSONL)A text format containing one valid JavaScript Object Notation value per line. Open glossary entry requests to /v1/responses each use a unique custom_id and complete asynchronously within a documented 24-hour window. It is not a real-time process. Outputs may be unordered, so the design must process both the output file and the error file. [official source 1]
Documented point: Strict JavaScript Object Notation (JSON)A text format for representing structured data as objects, arrays, numbers, strings, and other values. Open glossary entry Schema constrains response shape; explicit required fields and no unexpected properties support record contracts. Schema validity is not factual, policy, fairness, or business correctness; nullable unions represent intentional absence. [official source 1]
Documented point: A human-labelled holdout should gate the full batch and cover ambiguous, multi-issue, missing-context, and escalation-sensitive tickets. Choose the acceptance threshold as a risk decision, and use a local or independently maintained harness because the OpenAI API Evals platform is being retired. [official source 1 official source 2]
Documented point: As of the development date, this must be framed as an API workflow: GPT-5.5 is scheduled to retire from ChatGPT, ChatGPT Work and Codex, but not from the API. Revalidate lifecycle status and account access before each production run. [official source 1 official source 2]
Set the operating boundary before writing the classifier
This tutorial builds an offline, reviewable classifier for an authorised, read-only export of support tickets. The classifier assigns a controlled category, identifies defined review conditions, and records brief evidence for a person to verify. It does not send replies, close tickets, alter priorities, change customer accounts, trigger refunds, or update a help-desk system. Those actions remain outside the model workflow unless a separately authorised, human-controlled process is designed and approved.
OpenAI’s GPT-5.5 model reference, accessed on 30 September 2026, lists support for the Responses API, Batch API, and Structured Outputs. That combination supports the technical design used here: each ticket becomes a request to /v1/responses, the model must return an object matching a strict JSON schema, and non-urgent requests can be submitted as a batch. This is an API workflow, not a consumer ChatGPT procedure.
The product-surface distinction matters. OpenAI’s current documentation says GPT-5.5 is scheduled to retire from ChatGPT, ChatGPT Work, and Codex on 14 October 2026, while explicitly excluding the API from that notice. Revalidate the live model documentation, account access and lifecycle status before implementing or rerunning the pipeline. Do not infer the availability of a separately named GPT-5.5 Mini: the reviewed official sources document gpt-5.5 for this workflow.
The Batch API is suitable only when the work is not time-critical. OpenAI documents asynchronous processing with a 24-hour completion window, but that is a processing window rather than a promise that every request finishes at exactly 24 hours or earlier. Requests can fail or expire, and completed output need not preserve input order. The operating design must therefore retain the authorised input, assign a unique custom_id to every request, reconcile results by that identifier, inspect both output and error files, and define what may be reviewed or resubmitted.
Decision rule: use this pipeline for offline analysis or staging only. If a ticket requires an immediate safety, account-security, legal, regulatory or service response, route it through the existing human-led urgent process rather than waiting for a batch.
For support-operations prompt patterns for triage, reply quality assurance and escalation, see 25 ChatGPT-5.5 Prompts for Zendesk Support Operations: Ticket Triage, Reply QA, Escalation, and Knowledge Gaps.
For large-file batch processing in ChatGPT Work, see How to Process Large File Batches with ChatGPT Work: Complete Playbook for Document Analysis, Data Extraction, and Bulk Operations.
Write an explicit authorisation statement
Before extracting any data, create a short scope document approved by the system owner, support owner and relevant privacy or security reviewer. This document is an operational control, not a legal conclusion. It should identify the source system, permitted ticket population, approved fields, date range, processing purpose, recipients of results, retention period, deletion procedure and people authorised to review sensitive records.
A workable authorisation statement can be concise, but it must be specific enough to prevent scope drift. “Analyse support data” is inadequate because it does not distinguish classification from reply generation, customer profiling or account action. The following is a recommended template; replace the bracketed text with approved organisational details.
Purpose:
Classify an authorised, read-only export of [ticket population] into the
approved taxonomy for evaluation and human-reviewed routing analysis.
Permitted source:
[System and export owner]
Permitted date range:
[Start date] to [end date]
Permitted input fields:
[Approved field list]
Excluded records:
[Legal holds, restricted regions, employee cases, minors, health data,
payment data, credentials, or other organisation-specific exclusions]
Permitted outputs:
Category, review route, evidence excerpt or evidence field reference,
and machine-processing metadata.
Prohibited actions:
No automated reply, closure, priority change, refund, account change,
security action, disciplinary action, or customer-facing decision.
Review authority:
[Named role or team]
Retention and deletion:
[Approved duration, storage location, deletion owner, and evidence of deletion]
Assign an owner to approve each export rather than treating prior access to the help-desk system as continuing permission for model processing. Access rights and processing authority are different controls. A support analyst might be able to view a ticket while still lacking permission to export it, transfer it to another processing environment, or retain a derived dataset.
Define the unit of classification
Choose one stable unit before labelling. For this tutorial, one record represents one support ticket at the time of export. If a ticket contains a conversation, include only the approved messages required to identify its primary issue. Do not silently split messages into independent examples or combine unrelated tickets, because either change alters the meaning of the ground-truth label and makes evaluation results difficult to interpret.
Freeze a source timestamp or export version. Tickets can change while they remain open, so a label assigned to yesterday’s text may not describe today’s conversation. Store the source ticket identifier separately from the stable, opaque classifier identifier, and record the export timestamp. Use the opaque classifier identifier, not the source ticket identifier, as ticket_ref in model-facing records. The classifier identifier should not expose customer information; a generated opaque value such as ticket_000184 is preferable to an email address, subject line or account number.
If the source system permits repeated exports, define deduplication before sampling. A recommended rule is to treat the same source ticket at the same export version as one record, while a materially updated conversation becomes a new version linked to the original source identifier. This preserves change history without allowing duplicate examples to inflate holdout results.
Minimise ticket content before it reaches the API
Data minimisation starts with field selection, not prompt wording. Export only the content required to distinguish approved categories and review routes. A typical classifier may need a cleaned subject, selected conversation text, product area and channel. It usually does not need a customer’s full name, email address, telephone number, postal address, payment details, authentication data, internal account notes or unrelated historical tickets.
Do not rely on the model to ignore unnecessary personal or confidential information. Remove or transform fields before constructing JSONL. JSONL is a text format in which each line contains one complete JSON object; in the Batch workflow, each line will later represent an independent API request. Minimising the source table first prevents excluded data from being copied into request files, logs, test fixtures and debugging artefacts.
| Source field | Default treatment | Reason |
|---|---|---|
| Internal ticket identifier (ID)A value used to distinguish one record, task, source or object from another. Open glossary entry | Store in a restricted reconciliation table; replace in model input with an opaque classifier identifier | Supports traceability without exposing the operational identifier throughout the pipeline |
| Subject | Include only after redaction and relevance review | Often contains useful issue context but may also contain names, order numbers or credentials |
| Conversation text | Include the minimum approved messages; redact excluded content | Provides classification evidence while limiting unrelated history |
| Customer name and contact details | Exclude | Not normally required to classify the support issue |
| Passwords, access tokens or recovery codes | Exclude and route the source ticket to the established security process | Credentials must not become classifier inputs or examples |
| Payment-card or bank information | Exclude; follow the organisation’s approved incident and handling procedure | Classification does not justify processing high-risk financial data |
| Free-form internal notes | Exclude by default; include only specifically approved fields | Notes may contain sensitive judgements or information unrelated to the customer’s issue |
| Existing queue or tag | Keep outside the model input when it is the label being evaluated | Prevents answer leakage and allows an honest comparison with human labels |
Use deterministic redaction rules where possible and log only the type and count of transformations, not the removed value. For example, replace an email address with [EMAIL_REDACTED] and an apparent access token with [CREDENTIAL_REDACTED]. A redaction pattern can miss unusual formats or remove meaningful text, so sample the transformed records manually before submission. High-risk findings should stop processing for that record rather than merely replacing a substring and proceeding automatically.
Keep the mapping between opaque classifier identifiers and source ticket identifiers encrypted or otherwise protected under the organisation’s approved controls, with access limited to reviewers who need to return to the source. The model-facing dataset should not contain that mapping. Also avoid putting ticket content into filenames, command-line arguments, exception messages or general-purpose logs.
Apply a preflight exclusion gate
A recommended preflight gate should reject or quarantine records containing prohibited fields, empty text, unsupported encodings, excessive duplication, obvious secrets, or content types outside the authorised purpose. “Quarantine” means the record is withheld for authorised human assessment; it does not mean the model has determined that the record is dangerous or unlawful.
- Pass: the record contains only approved fields, has enough text for the task and does not trigger an exclusion rule.
- Quarantine: the record may contain credentials, payment data, account-security information, legal demands, threats, self-harm content, health information or another locally defined sensitive class.
- Reject: the record is outside the approved population, lacks a stable identifier, duplicates an existing export version or cannot be decoded safely.
Record the gate outcome and rule version for every source row. This gives reviewers a count of what was excluded and prevents silent loss between export and batch construction. It also permits the same transformation to be rerun when the taxonomy or model configuration changes.
Design a taxonomy that people can apply consistently
A taxonomy is the controlled set of labels and decision rules used by both human annotators and the classifier. Build it from the authorised business purpose, not from whatever tags happen to exist in historical data. Existing tags may reflect queue names, staffing arrangements, obsolete products or inconsistent agent habits rather than the customer’s actual issue.
Keep the first version small enough for annotators to distinguish reliably. Each category needs a stable code, plain-language name, inclusion rule, exclusion rule, examples and a tie-break rule. Category codes should remain stable even if display names change, because codes will be stored in holdout labels and machine output.
| Code | Category | Include when | Exclude or review when |
|---|---|---|---|
ACCESS_LOGIN |
Login or access difficulty | The primary issue is signing in, verification or access to an existing account | Route suspected takeover, credential exposure or identity dispute to human security review |
BILLING_QUERY |
Billing question | The customer asks for an explanation of an invoice, charge or subscription billing event | Do not infer fraud or approve refunds; disputed or sensitive payment cases require review |
PRODUCT_DEFECT |
Product malfunction | The customer reports expected functionality failing under described conditions | Exclude feature requests and cases with insufficient evidence to distinguish misuse from a defect |
FEATURE_REQUEST |
Requested capability | The customer asks for behaviour or functionality not currently described as available | Do not classify a documented malfunction as a feature request merely because a workaround exists |
HOW_TO |
Usage guidance | The customer asks how to perform an ordinary supported task | Exclude access failures, defects and policy disputes |
OTHER_REVIEW |
Unresolved or out-of-taxonomy | No approved category fits, context is missing, or issues are inseparable | Never force a more specific category merely to avoid this route |
The table is a tutorial example, not a universal support taxonomy. Replace it with labels approved by the relevant support and risk owners. In particular, do not create model-controlled categories that make legal findings, diagnose health conditions, determine fraud, judge a person’s intent, or decide whether an account should be restricted.
Separate category from review routing
A ticket can have an ordinary issue category and still require specialist review. Represent these as separate fields. For example, a suspected account takeover might retain the broad category ACCESS_LOGIN while setting review_route to ACCOUNT_SECURITY. This avoids expanding the taxonomy into a mixture of topic, severity, customer status and organisational workflow.
The review-route values below are a conservative starting point: STANDARD_REVIEW, ACCOUNT_SECURITY, LEGAL_OR_REGULATORY, SAFETY_SENSITIVE, PRIVACY_SENSITIVE and INSUFFICIENT_CONTEXT. These values should mean “send to an authorised person or existing process”, not “the model has confirmed the condition”.
If the schema includes any model-produced numeric self-assessment, treat it only as descriptive model output unless the organisation has separately tested its calibration. Strict Structured Outputs can require a number or enumeration to appear, but schema compliance does not establish that the number measures correctness. A safer initial design uses explicit observable reasons for review, such as missing context, conflicting issues, sensitive content or no matching category.
Write deterministic tie-break rules
Multi-issue tickets are unavoidable. Define whether the task returns one primary category or several categories before annotation. This tutorial uses one primary category so that acceptance metrics and downstream review remain interpretable. Annotators should choose the issue that blocks resolution of the ticket’s stated request; if no issue clearly dominates, they should use OTHER_REVIEW and record multiple_inseparable_issues as the reason.
Do not hide ambiguity by letting each annotator devise a separate priority rule. Add concrete rules to the taxonomy, such as: an inability to sign in takes precedence over a later how-to question when access blocks the requested task; a charge explanation remains BILLING_QUERY even if the customer also requests a refund, because the classifier is not authorised to approve that refund; and a security concern always adds specialist review regardless of the topical category.
Create an independently maintained human-labelled holdout
A holdout set is a frozen collection of representative records with ground-truth labels assigned through a documented human process. “Ground truth” here means the organisation’s approved reference decision for evaluation, not an assertion that every ticket has one objectively perfect interpretation. OpenAI’s evaluation guidance recommends representative examples and measured testing; the holdout should therefore resemble the intended ticket population while deliberately covering difficult cases.
Maintain the holdout locally or in another independently controlled evaluation harness. OpenAI documents that its API Evals platform is deprecated, becomes read-only on 31 October 2026 and shuts down on 30 November 2026. A durable gate should not depend on that retiring platform. Store the examples, labels, taxonomy version, annotation notes, evaluation code and results in your own controlled repository or data system.
Sample the holdout before prompt tuning and keep it separate from prompt examples, development fixtures and taxonomy drafting. If developers repeatedly inspect holdout failures and rewrite the prompt around individual cases, the set stops measuring generalisation and becomes another development set. Use a separate development set for iteration, then run the frozen holdout only at defined release gates.
Cover ordinary and difficult records
A representative holdout should reflect normal category prevalence, but prevalence alone is insufficient. Include deliberate slices for ambiguous wording, multiple issues, missing context, redacted content, very short messages, long conversations, contradictory statements, unsupported languages if they are in scope, out-of-taxonomy requests and escalation-sensitive cases. Report performance for these slices separately so a strong result on common easy tickets cannot conceal failures on consequential routes.
Use two qualified annotators for ambiguous or sensitive examples where practical, followed by adjudication by an authorised owner. Record the initial labels, disagreement reason, final label and adjudicator. Do not resolve disagreements by showing annotators the model’s answer; that can bias the reference label towards the system under evaluation.
The annotation record should include:
example_id: an opaque, stable identifier;taxonomy_version: the exact rule set used;primary_category: the approved category code;review_route: the approved human route;reason_codes: controlled ambiguity or sensitivity indicators;evidence_spans: minimal text offsets or approved excerpts supporting the decision;annotator_ids: internal pseudonymous reviewer identifiers;adjudication_status: agreed, adjudicated or unresolved;notes: concise explanation of the applied taxonomy rule.
Exclude unresolved examples from the headline category score until the policy owner determines the reference label, but retain them as a separate challenge set. Their presence can reveal that the taxonomy, rather than the model, needs revision.
Set acceptance criteria before seeing model results
Acceptance thresholds are risk decisions made by the organisation; OpenAI’s documentation does not supply a universal passing score for support classification. Define the gate before running GPT-5.5 so that stakeholders cannot lower standards after seeing an attractive aggregate result. The gate should combine structural validity, coverage, category performance, sensitive-route performance and manual error review.
| Gate | Decision rule | Failure response |
|---|---|---|
| Schema validity | Set an organisation-approved minimum and require every invalid or missing result to enter reconciliation | Inspect request, response and error records; do not count missing outputs as correct |
| Identifier integrity | Every expected custom_id appears exactly once across reconciled success and error states |
Stop scale-up and investigate duplicates, omissions or unknown identifiers |
| Primary category quality | Meet the pre-approved overall metric and the threshold for each designated high-risk slice | Revise taxonomy, prompt or model configuration using the development set, then rerun the frozen gate |
| Sensitive-route recall | Meet a separately approved conservative threshold for records that require specialist review | Do not release operational use; examine every false negative with the responsible owner |
| Error concentration | No unacceptable systematic failure by category, language, channel or other authorised evaluation slice | Restrict scope, improve data and rules, or reject deployment |
| Human review | Named reviewers approve the error analysis and intended operating boundary | No full-export run or downstream integration |
Use metrics that match the taxonomy. Overall accuracy can be reported for a single-label task, but it can mask weak performance on rare classes. Also calculate per-category precision and recall, a confusion matrix, review-route false negatives, unresolved-output counts and schema-failure counts. These are reader-run evaluation results, not promised GPT-5.5 performance.
Define hard-stop errors separately from metric thresholds. Examples include assigning an unknown category, failing to route an explicitly sensitive record for review, emitting source identifiers not present in the request, omitting expected records, or producing duplicate custom_id values. A single hard-stop error may justify rejecting a configuration even when its aggregate score passes.
Strict Structured Outputs helps enforce required fields, permitted enumerations and additionalProperties: false. It does not guarantee that the chosen category is true, fair, safe or appropriate. Nullable unions should represent intentional absence, such as no evidence excerpt being available; they should not be used to disguise processing failure. Human review remains mandatory for escalation, legal, account-security, privacy-sensitive and other consequential routes.
Prepare a reproducible project layout
Separate raw authorised material, minimised records, annotations, API requests, returned files and evaluation reports. The following recommended layout keeps immutable evidence apart from generated artefacts. Apply organisational access controls rather than assuming a directory name provides security.
ticket-classifier/
├── README.md
├── config/
│ ├── taxonomy.v1.json
│ ├── output-schema.v1.json
│ ├── acceptance-criteria.v1.json
│ └── run-config.example.json
├── data/
│ ├── restricted/
│ │ ├── source-export/
│ │ └── id-map/
│ ├── minimised/
│ │ ├── development.jsonl
│ │ ├── holdout.jsonl
│ │ └── full-export.jsonl
│ └── labels/
│ ├── development-labels.jsonl
│ ├── holdout-labels.jsonl
│ └── adjudication-log.jsonl
├── batch/
│ ├── holdout/
│ │ ├── requests.jsonl
│ │ ├── manifest.json
│ │ ├── output.jsonl
│ │ └── errors.jsonl
│ └── production/
│ ├── requests.jsonl
│ ├── manifest.json
│ ├── output.jsonl
│ └── errors.jsonl
├── reports/
│ ├── holdout-metrics.json
│ ├── confusion-matrix.csv
│ ├── reconciliation.csv
│ └── human-review-signoff.md
├── scripts/
│ ├── minimise.py
│ ├── build_batch.py
│ ├── validate_jsonl.py
│ ├── reconcile.py
│ └── evaluate.py
└── tests/
├── fixtures/
└── test_contracts.py
Do not commit restricted exports, identifier maps, live ticket text, API credentials or returned sensitive content to ordinary source control. The repository can contain schemas, synthetic fixtures, scripts and redacted examples. Store operational data in an approved restricted location, and make paths configurable so developers do not copy live records into local test directories.
The run manifest should record the model identifier, model snapshot when appropriate, reasoning-effort setting, prompt version, taxonomy version, output-schema version, transformation-code revision, source-export checksum, request-file checksum, creation time and operator. OpenAI recommends measured reasoning tuning rather than assuming a higher setting is always better. Any change to model snapshot, prompt, reasoning effort, taxonomy, schema or preprocessing should trigger the relevant holdout gate again.
Exit criteria for this preparation stage
Do not construct the full Batch request file until all preparation checks pass. The authorised scope must be signed off; excluded fields must be removed; the taxonomy must have stable codes and tie-break rules; the holdout must be frozen and independently labelled; acceptance criteria must be approved; and the project must have protected locations for source data, ID maps, requests, outputs, errors and reports.
- The workflow is explicitly read-only and offline.
- Every record has an opaque, stable identifier and a traceable export version.
- Sensitive records have a human route that does not depend on successful batch completion.
- Model inputs exclude unnecessary personal data, credentials and restricted fields.
- Ground-truth labels were created without exposing annotators to model predictions.
- The holdout includes ambiguous, multi-issue, missing-context and escalation-sensitive examples.
- Acceptance thresholds and hard-stop conditions were fixed before evaluation.
- Operational files will be reconciled by
custom_id, never by line position. - No classifier output is authorised to reply, close, reprioritise or act on a customer account.
Turn the classification contract into strict Batch requests
OpenAI’s GPT-5.5 model reference, accessed on 30 September 2026, lists support for the Responses API, Batch API and Structured Outputs. This section uses those capabilities to produce JSONL requests whose individual target is /v1/responses. The workflow is API-specific: GPT-5.5’s scheduled retirement from ChatGPT, ChatGPT Work and Codex on 14 October 2026 explicitly excludes the API, although model availability and account access should still be rechecked before a production run.
The build starts from an authorised, minimised ticket export and produces an immutable JSONL input file. It does not reply to customers, close tickets, alter priority, change account state or write classifications into a help-desk system. Those actions remain outside this classifier’s authority and require a separately approved, human-led workflow.
OpenAI documents Batch as asynchronous processing with a 24-hour completion window. Treat that as a service window rather than a promise that every request will finish at a particular time. A batch can include failed or expired work, and its output order need not match its input order. Consequently, every request needs a unique custom_id, and the local input manifest must remain available until output and error files have been reconciled.
Use a small, auditable source-record format
Local format: keep the source export separate from the API request file. The source record should contain only the fields approved for classification, while the manifest should preserve the mapping between an internal ticket reference and the generated custom_id. This separation makes it possible to regenerate requests without repeatedly handling a larger operational export.
The example below assumes that the preflight process from the previous section has already removed unauthorised records and unnecessary personal data. It deliberately excludes customer names, email addresses, account identifiers, attachments, internal access tokens and free-form agent notes that are not required to apply the taxonomy.
{
"ticket_ref": "TICKET-EXAMPLE-0001",
"subject": "Cannot reset password after changing phone",
"body": "The reset flow asks for a code sent to an old phone number.",
"channel": "web",
"language": "en",
"created_at": "2026-09-28T10:15:00Z"
}
This is an illustrative local record, not a required OpenAI format. Replace the fields with the minimum approved by the organisation’s data owner. If a ticket contains payment data, legal threats, account-recovery evidence, health information or other sensitive content, route it through the preflight exclusion policy rather than assuming that schema-constrained output makes the input appropriate to process.
Define one strict JSON Schema for every result
Structured Outputs applies a JSON Schema to the model’s response. OpenAI’s guide documents strict mode, explicit required fields and additionalProperties: false. These controls give downstream code a predictable shape, but they do not establish that a category is true, fair, safe or suitable for an operational decision.
Schema design: return a taxonomy label, a review route, a concise evidence excerpt and an explanation of the classification. The evidence field should quote or closely identify content from the supplied ticket rather than introduce external facts. A nullable field represents intentional absence; omitting a required property does not.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"ticket_ref": {
"type": "string",
"description": "The ticket reference copied exactly from the input."
},
"category": {
"type": "string",
"enum": [
"ACCESS_LOGIN",
"BILLING_QUERY",
"PRODUCT_DEFECT",
"FEATURE_REQUEST",
"HOW_TO",
"OTHER_REVIEW"
]
},
"review_route": {
"type": "string",
"enum": [
"STANDARD_REVIEW",
"ACCOUNT_SECURITY",
"LEGAL_OR_REGULATORY",
"SAFETY_SENSITIVE",
"PRIVACY_SENSITIVE",
"INSUFFICIENT_CONTEXT"
]
},
"secondary_category": {
"type": ["string", "null"],
"enum": [
"ACCESS_LOGIN",
"BILLING_QUERY",
"PRODUCT_DEFECT",
"FEATURE_REQUEST",
"HOW_TO",
"OTHER_REVIEW",
null
]
},
"evidence": {
"type": "string",
"description": "A short excerpt or precise reference to the supplied ticket text."
},
"rationale": {
"type": "string",
"description": "A concise application of the supplied taxonomy and tie-break rules."
},
"needs_human_review": {
"type": "boolean"
},
"review_reasons": {
"type": "array",
"items": {
"type": "string",
"enum": [
"MULTIPLE_INSEPARABLE_ISSUES",
"MISSING_CONTEXT",
"SENSITIVE_CONTENT",
"ACCOUNT_SECURITY",
"LEGAL_OR_REGULATORY",
"NO_MATCHING_CATEGORY"
]
}
}
},
"required": [
"ticket_ref",
"category",
"review_route",
"secondary_category",
"evidence",
"rationale",
"needs_human_review",
"review_reasons"
],
"additionalProperties": false
}
The schema requires every field, including secondary_category. The nullable union allows the model to return null when no second category applies, while preserving a stable record shape. Do not use an absent property to mean “none”, “unknown” and “not evaluated” interchangeably.
The schema uses the approved taxonomy codes and route values defined above; replace those values only through the documented taxonomy-change process.
The needs_human_review boolean is a routing instruction, not a calibrated probability. It must not be presented as confidence. The enumerated review_reasons make the route inspectable, but the model can still select an inappropriate reason. Sensitive and consequential categories should therefore be forced into review through deterministic local rules as well as prompt instructions.
For structured-output validation and reliable JSON contracts, see How to Use ChatGPT Structured Outputs for Reliable JSON: Complete API Tutorial with Schema Validation and Error Handling.
Keep the prompt compact and preserve the contract
OpenAI’s GPT-5.5 guidance recommends representative examples and small prompt changes that preserve the task contract. For this classifier, the prompt should state the taxonomy, tie-break rules, review boundary and prohibition on unsupported inference. Avoid adding many loosely related instructions between holdout runs because doing so makes regressions difficult to attribute.
Prompt template: separate stable policy instructions from ticket data. The ticket should be serialised as data rather than interpolated into an instruction sentence. This reduces ambiguity about which text defines the task, although it does not eliminate malicious or misleading content inside a ticket.
SYSTEM_INSTRUCTIONS = """
You classify one support ticket using only the supplied ticket data.
Apply these rules:
1. Select one primary category from the allowed schema values.
2. Use a secondary category only when a distinct second issue is explicit.
3. Copy ticket_ref exactly.
4. Base evidence and rationale only on the supplied ticket.
5. Do not infer identity, entitlement, account ownership, legal status, fraud,
security compromise, or customer intent beyond the supplied text.
6. Set needs_human_review to true for ambiguity, multiple issues, missing
context, sensitive content, account-security matters, legal or regulatory
matters, or a taxonomy gap.
7. Set review_route to ACCOUNT_SECURITY, LEGAL_OR_REGULATORY,
SAFETY_SENSITIVE or PRIVACY_SENSITIVE as applicable.
8. Set review_route to INSUFFICIENT_CONTEXT when the requested classification
cannot be supported from the supplied text.
9. Classification is advisory. Do not recommend or perform replies, closure,
priority changes, refunds, access changes, or other customer-account action.
"""
If the approved taxonomy contains organisation-specific definitions, insert them into this stable instruction block and version the resulting text. Record the prompt version beside the model identifier and schema version so that a holdout result can be traced to the exact configuration that produced it.
For OpenAI Agents API workflow design, recovery and acceptance testing, see OpenAI Agents API workflow evaluation prompts.
Generate unique identifiers without exposing operational keys
The Batch guide requires a unique custom_id for each request. It is also the reliable join key when outputs arrive in a different order. Do not use a list position, output line number or ticket subject as the join key.
Identifier rule: derive a non-secret digest from a run identifier and the internal ticket reference. Prefixing the digest with the run identifier helps operators distinguish batches while avoiding direct exposure of the original ticket reference in the Batch envelope. The mapping remains in a protected local manifest.
from hashlib import sha256
def make_custom_id(run_id: str, ticket_ref: str) -> str:
material = f"{run_id}\x1f{ticket_ref}".encode("utf-8")
digest = sha256(material).hexdigest()[:24]
return f"{run_id}-{digest}"
A digest is not encryption. If ticket references are predictable, a party holding candidate ticket references may be able to hash them and compare the results with the identifiers. Keep the manifest under the same access controls as the source export, and use a random run identifier rather than embedding a customer, queue or account name.
Before writing JSONL, assert uniqueness locally. A duplicate should stop generation rather than be silently renamed, because it can indicate duplicate source records, unstable identifiers or a bug in the manifest logic.
def assert_unique_custom_ids(rows: list[dict]) -> None:
seen: set[str] = set()
duplicates: list[str] = []
for row in rows:
custom_id = row["custom_id"]
if custom_id in seen:
duplicates.append(custom_id)
seen.add(custom_id)
if duplicates:
raise ValueError(
"Duplicate custom_id values detected: "
+ ", ".join(sorted(set(duplicates)))
)
Build each /v1/responses request
Each JSONL line is a Batch request envelope. Its url is /v1/responses, while the request body contains the GPT-5.5 model, input and strict response format. The model reference documents Structured Outputs and Batch support for gpt-5.5. If the organisation pins a documented model snapshot, store that exact identifier in configuration and rerun the holdout before changing it.
import json
from typing import Any
MODEL_ID = "gpt-5.5"
SCHEMA_VERSION = "ticket-classification-v1"
PROMPT_VERSION = "classifier-instructions-v1"
CLASSIFICATION_SCHEMA: dict[str, Any] = {
"type": "object",
"properties": {
"ticket_ref": {"type": "string"},
"category": {
"type": "string",
"enum": [
"ACCESS_LOGIN",
"BILLING_QUERY",
"PRODUCT_DEFECT",
"FEATURE_REQUEST",
"HOW_TO",
"OTHER_REVIEW",
],
},
"review_route": {
"type": "string",
"enum": [
"STANDARD_REVIEW",
"ACCOUNT_SECURITY",
"LEGAL_OR_REGULATORY",
"SAFETY_SENSITIVE",
"PRIVACY_SENSITIVE",
"INSUFFICIENT_CONTEXT",
],
},
"secondary_category": {
"type": ["string", "null"],
"enum": [
"ACCESS_LOGIN",
"BILLING_QUERY",
"PRODUCT_DEFECT",
"FEATURE_REQUEST",
"HOW_TO",
"OTHER_REVIEW",
None,
],
},
"evidence": {"type": "string"},
"rationale": {"type": "string"},
"needs_human_review": {"type": "boolean"},
"review_reasons": {
"type": "array",
"items": {
"type": "string",
"enum": [
"MULTIPLE_INSEPARABLE_ISSUES",
"MISSING_CONTEXT",
"SENSITIVE_CONTENT",
"ACCOUNT_SECURITY",
"LEGAL_OR_REGULATORY",
"NO_MATCHING_CATEGORY",
],
},
},
},
"required": [
"ticket_ref",
"category",
"review_route",
"secondary_category",
"evidence",
"rationale",
"needs_human_review",
"review_reasons",
],
"additionalProperties": False,
}
def build_batch_request(
custom_id: str,
ticket: dict[str, Any],
) -> dict[str, Any]:
ticket_payload = {
"ticket_ref": ticket["ticket_ref"],
"subject": ticket.get("subject", ""),
"body": ticket.get("body", ""),
"channel": ticket.get("channel"),
"language": ticket.get("language"),
"created_at": ticket.get("created_at"),
}
return {
"custom_id": custom_id,
"method": "POST",
"url": "/v1/responses",
"body": {
"model": MODEL_ID,
"input": [
{
"role": "system",
"content": [
{
"type": "input_text",
"text": SYSTEM_INSTRUCTIONS,
}
],
},
{
"role": "user",
"content": [
{
"type": "input_text",
"text": json.dumps(
ticket_payload,
ensure_ascii=False,
separators=(",", ":"),
),
}
],
},
],
"text": {
"format": {
"type": "json_schema",
"name": "ticket_classification",
"strict": True,
"schema": CLASSIFICATION_SCHEMA,
}
},
},
}
This builder copies the approved fields explicitly rather than forwarding the entire source dictionary. That allowlist prevents a later export change from silently adding a new field to the API payload. Treat the allowlist as a data-minimisation control and review it whenever the export format changes.
The system instruction tells the model to copy ticket_ref, but that returned value is not the authoritative join key. Reconciliation must use the envelope’s custom_id and compare the returned ticket_ref with the protected manifest. A mismatch should enter the error queue rather than being accepted or joined to another record.

Write JSONL and an independent manifest atomically
JSONL requires one complete JSON object per line. Do not pretty-print the request file, concatenate arrays or allow embedded newline characters outside JSON string escaping. Use a temporary file, flush it to disk and rename it only after all records pass validation.
import json
import os
from pathlib import Path
from uuid import uuid4
def write_batch_files(
tickets: list[dict[str, Any]],
output_dir: Path,
) -> tuple[Path, Path, str]:
output_dir.mkdir(parents=True, exist_ok=True)
run_id = f"run-{uuid4().hex[:12]}"
requests_path = output_dir / f"{run_id}.requests.jsonl"
manifest_path = output_dir / f"{run_id}.manifest.jsonl"
requests_tmp = requests_path.with_suffix(".jsonl.tmp")
manifest_tmp = manifest_path.with_suffix(".jsonl.tmp")
request_rows: list[dict[str, Any]] = []
manifest_rows: list[dict[str, Any]] = []
for ticket in tickets:
ticket_ref = str(ticket["ticket_ref"])
custom_id = make_custom_id(run_id, ticket_ref)
request_rows.append(
build_batch_request(
custom_id=custom_id,
ticket=ticket,
)
)
manifest_rows.append(
{
"custom_id": custom_id,
"ticket_ref": ticket_ref,
"run_id": run_id,
"model": MODEL_ID,
"schema_version": SCHEMA_VERSION,
"prompt_version": PROMPT_VERSION,
"source_digest": sha256(
json.dumps(
ticket,
ensure_ascii=False,
sort_keys=True,
separators=(",", ":"),
).encode("utf-8")
).hexdigest(),
}
)
assert_unique_custom_ids(request_rows)
try:
with requests_tmp.open("x", encoding="utf-8") as request_file:
for row in request_rows:
request_file.write(
json.dumps(row, ensure_ascii=False, separators=(",", ":"))
)
request_file.write("\n")
request_file.flush()
os.fsync(request_file.fileno())
with manifest_tmp.open("x", encoding="utf-8") as manifest_file:
for row in manifest_rows:
manifest_file.write(
json.dumps(row, ensure_ascii=False, separators=(",", ":"))
)
manifest_file.write("\n")
manifest_file.flush()
os.fsync(manifest_file.fileno())
requests_tmp.replace(requests_path)
manifest_tmp.replace(manifest_path)
except Exception:
requests_tmp.unlink(missing_ok=True)
manifest_tmp.unlink(missing_ok=True)
raise
return requests_path, manifest_path, run_id
The manifest’s source_digest detects accidental local changes; it is not a signature or security guarantee. Store the source file, request file, manifest, schema and prompt version as a controlled run bundle. Restrict access and apply the organisation’s approved retention schedule rather than keeping ticket text indefinitely for convenience.
Validate locally before uploading
Strict Structured Outputs constrains generated responses, but it does not excuse malformed Batch input. Perform inexpensive local checks before any network call: parse every line independently, verify unique identifiers, require the expected method and Uniform Resource Locator (URL)The address used to identify and access a resource on the web. Open glossary entry, check the configured model, and confirm that strict mode and additionalProperties: false are present.
def validate_request_file(path: Path) -> int:
seen: set[str] = set()
count = 0
with path.open("r", encoding="utf-8") as handle:
for line_number, line in enumerate(handle, start=1):
if not line.strip():
raise ValueError(f"Blank line at {line_number}")
row = json.loads(line)
custom_id = row["custom_id"]
if custom_id in seen:
raise ValueError(
f"Duplicate custom_id at line {line_number}: {custom_id}"
)
seen.add(custom_id)
if row["method"] != "POST":
raise ValueError(f"Unexpected method at line {line_number}")
if row["url"] != "/v1/responses":
raise ValueError(f"Unexpected URL at line {line_number}")
body = row["body"]
if body["model"] != MODEL_ID:
raise ValueError(f"Unexpected model at line {line_number}")
response_format = body["text"]["format"]
if response_format.get("type") != "json_schema":
raise ValueError(f"Missing JSON Schema at line {line_number}")
if response_format.get("strict") is not True:
raise ValueError(f"Strict mode disabled at line {line_number}")
if response_format["schema"].get("additionalProperties") is not False:
raise ValueError(
f"Additional properties not disabled at line {line_number}"
)
count += 1
if count == 0:
raise ValueError("Request file contains no records")
return count
Additional check: inspect a redacted sample from the generated file and compare it with the approved field allowlist. Automated tests can confirm structure, but a human reviewer is better placed to notice that an apparently harmless field contains account-security notes or unnecessary personal information.
Load credentials from the environment, not the project
Use an environment variable for the API credential and keep local secret files outside version control. Never place a credential in Python source, JSONL, a notebook, a shell-history example, a manifest or a committed configuration file.
import os
from openai import OpenAI
api_key = os.environ.get("OPENAI_API_KEY")
if not api_key:
raise RuntimeError(
"OPENAI_API_KEY is not set in the current process environment"
)
client = OpenAI(api_key=api_key)
The variable name above is a local naming convention, not a credential. Populate it through the organisation’s approved secret manager or a temporary process environment. If a local .env file is permitted, exclude it from version control, limit filesystem permissions and avoid loading it in shared build logs. Secret handling remains an organisational security decision; this example does not establish that a workstation or deployment environment is secure.
Upload once, record the file identifier, then create the Batch job
OpenAI’s Batch guide documents uploading the JSONL file for Batch processing and creating a job whose endpoint is /v1/responses with the documented 24-hour completion window. The following code uses the official Python client pattern and immediately records returned identifiers in a local submission receipt.
from datetime import datetime, timezone
def upload_and_create_batch(
client: OpenAI,
requests_path: Path,
receipt_path: Path,
) -> dict[str, Any]:
if receipt_path.exists():
raise FileExistsError(
f"Submission receipt already exists: {receipt_path}"
)
with requests_path.open("rb") as source:
uploaded = client.files.create(
file=source,
purpose="batch",
)
batch = client.batches.create(
input_file_id=uploaded.id,
endpoint="/v1/responses",
completion_window="24h",
metadata={
"workflow": "ticket-classification",
"schema_version": SCHEMA_VERSION,
"prompt_version": PROMPT_VERSION,
},
)
receipt = {
"created_at": datetime.now(timezone.utc).isoformat(),
"request_file": requests_path.name,
"input_file_id": uploaded.id,
"batch_id": batch.id,
"endpoint": "/v1/responses",
"model": MODEL_ID,
"schema_version": SCHEMA_VERSION,
"prompt_version": PROMPT_VERSION,
}
temporary = receipt_path.with_suffix(".json.tmp")
temporary.write_text(
json.dumps(receipt, indent=2, sort_keys=True),
encoding="utf-8",
)
temporary.replace(receipt_path)
return receipt
The existence check prevents an operator from accidentally submitting the same local file twice. It is a local safeguard, not a claim of server-side idempotency. If a process stops after the server accepts an upload or Batch creation but before the receipt is written, do not automatically create another job. First inspect the account’s Batch records and local logs, identify whether a job already exists, and have an operator decide whether resubmission is necessary.
Retry transport failures without duplicating accepted work
Retry policy: retry only failures that are plausibly transient, use bounded exponential delay with random jitter, and separate upload retry from Batch-creation retry. Authentication failures, permission errors, invalid JSONL, unsupported configuration and schema defects require correction rather than repeated calls.
import random
import time
from collections.abc import Callable
from typing import TypeVar
T = TypeVar("T")
class RetryableSubmissionError(Exception):
pass
def call_with_bounded_retry(
operation: Callable[[], T],
attempts: int = 4,
base_delay_seconds: float = 1.0,
) -> T:
last_error: Exception | None = None
for attempt in range(attempts):
try:
return operation()
except RetryableSubmissionError as exc:
last_error = exc
if attempt == attempts - 1:
break
delay = base_delay_seconds * (2 ** attempt)
delay += random.uniform(0, base_delay_seconds)
time.sleep(delay)
assert last_error is not None
raise last_error
This helper intentionally does not guess which client exceptions are retryable. Map only the transport and service conditions approved by the application’s error-handling policy. Do not catch every exception and retry it: doing so can repeatedly submit an invalid file or conceal a programming error.
Batch creation has an additional point of uncertainty. A connection can fail after the server has accepted a request but before the client receives the response. Because an unrecorded retry could create duplicate work, write an “attempt started” journal entry before creation, capture any returned identifier immediately, and require an account-side lookup or operator decision when acceptance is uncertain.
Operational rule: retries may recover communication failures; they must not convert uncertainty about an accepted Batch job into an automatic duplicate submission.
Submit the holdout before the full export
The first uploaded request file should contain the independently maintained, human-labelled holdout rather than the full ticket export. It should include ordinary cases as well as ambiguous, multi-issue, missing-context and escalation-sensitive records. The same schema, prompt, model setting and local validation code must be used for both the holdout and any later scale-up run.
The API Evals platform is being retired, becoming read-only on 31 October 2026 and shutting down on 30 November 2026. Keep the holdout, labels, scoring code and acceptance decision in an independently maintained local harness rather than making this workflow depend on that platform.
Do not weaken the strict schema or remove difficult records to obtain a passing result. A failed gate should produce a versioned change to the taxonomy, prompt, schema or preprocessing rules, followed by a fresh run. If the model identifier or snapshot changes, rerun the same gate before generating a full-export Batch file.
Build a submission checklist
- Authorisation: confirm that every source record is permitted for this API workflow and that unnecessary fields have been removed.
- Taxonomy: freeze the category definitions, tie-break rules and sensitive-review policy for the run.
- Schema: require every property, use nullable unions deliberately and set
additionalPropertiestofalse. - Identifiers: generate one unique
custom_idper request and preserve its protected mapping to the source ticket. - Payload: allowlist fields explicitly and target
/v1/responsesin each JSONL request. - Configuration: record the model identifier, prompt version and schema version; rerun evaluation after changes.
- Local validation: parse every JSONL line, reject duplicates and inspect a redacted human-readable sample.
- Secrets: load the API credential from an approved secret path and keep it out of source files, artefacts and logs.
- Submission safety: create an attempt journal and receipt so an uncertain network result does not trigger an automatic duplicate Batch.
- Scale gate: submit only the holdout until human reviewers have evaluated classifications and review routes against the predeclared acceptance rules.
At the end of this build stage, the expected artefacts are a validated holdout JSONL file, a protected manifest, a versioned schema, a versioned instruction block and a submission receipt containing the uploaded file and Batch identifiers. Preserve all five: the next stage must reconcile unordered results and error records by custom_id, not by position or assumption.
Gate the production run on holdout evidence
Run the human-labelled holdout through the same model identifier or pinned model snapshot, prompt, JSON Schema, reasoning configuration and Batch request builder intended for the full export. OpenAI’s GPT-5.5 documentation lists support for the Responses API, Batch API and Structured Outputs, but that capability statement does not establish that a particular taxonomy will be classified correctly. The holdout is the evidence for deciding whether this configuration is suitable for this dataset.
Use the independently maintained holdout; because it is stored outside the API Evals platform, it remains the evaluation record after that platform is retired. A local harness consisting of versioned labels, request manifests, downloaded responses and scoring code remains usable across that transition and makes later model or prompt comparisons reproducible.
The holdout gate should answer three separate questions: whether every successful model response conforms to the expected contract, whether its labels agree sufficiently with human ground truth, and whether the review-routing rules catch sensitive or ambiguous cases. Do not combine these into one headline score. A classifier can have high overall agreement while performing poorly on a rare escalation class that carries greater operational risk.
Freeze the evaluated configuration
Before downloading results, write an immutable run record. The record should identify the holdout version, taxonomy version, prompt digest, schema digest, request JSONL digest, manifest digest, model identifier or snapshot, reasoning setting, Batch identifier and submission time. This prevents a later prompt edit or taxonomy change from being mistaken for the configuration that produced the score.
{
"run_id": "holdout-2026-09-30-a",
"dataset_version": "holdout-v3",
"taxonomy_version": "support-taxonomy-v4",
"model": "gpt-5.5",
"prompt_sha256": "<computed digest>",
"schema_sha256": "<computed digest>",
"request_jsonl_sha256": "<computed digest>",
"manifest_sha256": "<computed digest>",
"batch_id": "<returned batch identifier>",
"submitted_at": "2026-09-30T14:00:00Z",
"purpose": "offline holdout evaluation only"
}
If the model reference offers an appropriate snapshot for the account and use case, recording that snapshot improves reproducibility. Whether using a snapshot or a moving alias, rerun the holdout after changing the model, prompt, schema, reasoning effort, taxonomy or input-minimisation rules. OpenAI’s GPT-5.5 guidance recommends representative examples and measured tuning rather than assuming that a configuration change is harmless.
Submit the holdout as an asynchronous Batch job
The Batch API is appropriate here because the classification is non-urgent and offline. OpenAI documents JSONL input, a unique custom_id for each request, use of /v1/responses, and an asynchronous completion window documented as 24 hours. That is a completion window, not a promise that every request will complete successfully or at an exact time.
The following Python example is a recommended implementation pattern. It assumes the validated holdout JSONL from the preceding build stage already exists. It deliberately stores identifiers and exits instead of waiting synchronously for classification results.
import json
import os
from datetime import datetime, timezone
from pathlib import Path
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
input_path = Path("runs/holdout-2026-09-30-a/requests.jsonl")
run_path = Path("runs/holdout-2026-09-30-a/run.json")
with input_path.open("rb") as source:
uploaded = client.files.create(
file=source,
purpose="batch",
)
batch = client.batches.create(
input_file_id=uploaded.id,
endpoint="/v1/responses",
completion_window="24h",
metadata={
"run_id": "holdout-2026-09-30-a",
"dataset": "holdout-v3",
},
)
run_record = {
"run_id": "holdout-2026-09-30-a",
"input_file_id": uploaded.id,
"batch_id": batch.id,
"status": batch.status,
"submitted_at": datetime.now(timezone.utc).isoformat(),
}
run_path.write_text(
json.dumps(run_record, indent=2, sort_keys=True),
encoding="utf-8",
)
print(json.dumps(run_record, indent=2))
Do not submit the full ticket export merely because the holdout upload was accepted. Acceptance confirms that a Batch job was created; it does not establish semantic quality. The scale-up decision comes only after terminal-state handling, output reconciliation, quantitative scoring and manual review.
For GPT-5.5 migration planning across evaluations, permissions, cost, rollback and sign-off, see GPT-5.5 to GPT-5.6 and Codex Migration Playbook: Inventory, Representative Evals, Tool Permissions, Cost, Rollback, and Sign-Off.
Poll conservatively and preserve every observed state
Polling should tolerate a job remaining asynchronous for an extended period. Use increasing intervals, add a maximum local wait if required by the scheduler, and record every status transition. A local timeout means “stop polling and inspect later”; it must not be translated into “the Batch failed” unless the API reports a failed or expired terminal state.
import json
import random
import time
from datetime import datetime, timezone
from pathlib import Path
from openai import OpenAI
client = OpenAI()
run_dir = Path("runs/holdout-2026-09-30-a")
run = json.loads((run_dir / "run.json").read_text(encoding="utf-8"))
terminal_states = {"completed", "failed", "expired", "cancelled"}
delay_seconds = 15
maximum_delay_seconds = 300
while True:
batch = client.batches.retrieve(run["batch_id"])
event = {
"observed_at": datetime.now(timezone.utc).isoformat(),
"batch_id": batch.id,
"status": batch.status,
"output_file_id": getattr(batch, "output_file_id", None),
"error_file_id": getattr(batch, "error_file_id", None),
}
with (run_dir / "status-events.jsonl").open("a", encoding="utf-8") as log:
log.write(json.dumps(event, sort_keys=True) + "\n")
print(event)
if batch.status in terminal_states:
break
jitter = random.uniform(0, delay_seconds * 0.1)
time.sleep(delay_seconds + jitter)
delay_seconds = min(delay_seconds * 2, maximum_delay_seconds)
Inspect any Batch-level error information returned by the client as well as the request-level error file. A failed Batch can indicate a job-level problem, while a completed Batch can still require scrutiny of individual request outcomes. Treat completed as permission to reconcile files, not as proof that every source record received a usable classification.

Download both the output and error artefacts
OpenAI’s Batch guide documents output and error files. Download every file identifier that is present, save the bytes without modification, and compute a digest before parsing. The raw artefacts are the audit source for later reconciliation; a cleaned spreadsheet alone cannot show whether a parser dropped a line or transformed a value.
import hashlib
import json
from pathlib import Path
from openai import OpenAI
client = OpenAI()
run_dir = Path("runs/holdout-2026-09-30-a")
run = json.loads((run_dir / "run.json").read_text(encoding="utf-8"))
batch = client.batches.retrieve(run["batch_id"])
def download_file(file_id: str, destination: Path) -> dict:
content = client.files.content(file_id)
raw = content.read()
destination.write_bytes(raw)
return {
"file_id": file_id,
"path": str(destination),
"bytes": len(raw),
"sha256": hashlib.sha256(raw).hexdigest(),
}
downloads = []
if getattr(batch, "output_file_id", None):
downloads.append(
download_file(batch.output_file_id, run_dir / "raw-output.jsonl")
)
if getattr(batch, "error_file_id", None):
downloads.append(
download_file(batch.error_file_id, run_dir / "raw-errors.jsonl")
)
(run_dir / "downloads.json").write_text(
json.dumps(downloads, indent=2, sort_keys=True),
encoding="utf-8",
)
Use restrictive filesystem permissions appropriate to the ticket data and organisational policy. The downloaded files may contain minimised ticket text, model-generated summaries or sensitive classifications. Do not place them in a public repository, general analytics bucket or shared spreadsheet merely because the original export was authorised.
Reconcile by custom_id, never by line position
OpenAI explicitly warns that Batch output may not preserve input order. Joining the first output line to the first source ticket can silently assign a valid classification to the wrong person or case. Reconciliation must therefore use the unique custom_id stored in both the request and the independently generated manifest.
The reconciler should reject duplicate identifiers, unknown identifiers and missing identifiers before calculating quality metrics. A syntactically valid response associated with an unrecognised custom_id is not safe to guess into place. Quarantine it and investigate the request-generation or file-selection process.
import json
from pathlib import Path
run_dir = Path("runs/holdout-2026-09-30-a")
def read_jsonl(path: Path):
if not path.exists():
return []
rows = []
with path.open(encoding="utf-8") as handle:
for line_number, line in enumerate(handle, start=1):
if not line.strip():
continue
try:
rows.append((line_number, json.loads(line)))
except json.JSONDecodeError as exc:
raise ValueError(
f"{path}: invalid JSON on line {line_number}: {exc}"
) from exc
return rows
manifest_rows = {
row["custom_id"]: row
for _, row in read_jsonl(run_dir / "manifest.jsonl")
}
if len(manifest_rows) != len(read_jsonl(run_dir / "manifest.jsonl")):
raise ValueError("Manifest contains duplicate custom_id values")
outcomes = {}
for line_number, row in read_jsonl(run_dir / "raw-output.jsonl"):
custom_id = row.get("custom_id")
if not custom_id:
raise ValueError(f"Output line {line_number} has no custom_id")
if custom_id not in manifest_rows:
raise ValueError(f"Unknown output custom_id: {custom_id}")
if custom_id in outcomes:
raise ValueError(f"Duplicate outcome for custom_id: {custom_id}")
outcomes[custom_id] = {
"source": "output",
"line_number": line_number,
"row": row,
}
for line_number, row in read_jsonl(run_dir / "raw-errors.jsonl"):
custom_id = row.get("custom_id")
if not custom_id:
raise ValueError(f"Error line {line_number} has no custom_id")
if custom_id not in manifest_rows:
raise ValueError(f"Unknown error custom_id: {custom_id}")
if custom_id in outcomes:
raise ValueError(f"Multiple outcomes for custom_id: {custom_id}")
outcomes[custom_id] = {
"source": "error",
"line_number": line_number,
"row": row,
}
missing = sorted(set(manifest_rows) - set(outcomes))
print({
"manifest_records": len(manifest_rows),
"reconciled_outcomes": len(outcomes),
"missing_outcomes": len(missing),
})
A missing outcome is a processing state, not a negative label and not permission to infer the result from neighbouring records. Keep missing records unresolved until the Batch status, output file, error file and retained input have been inspected. If they remain unresolved, route them either to a carefully controlled resubmission or to human classification.
Extract and revalidate the structured result
Strict Structured Outputs constrains the response to the supplied JSON Schema when the request succeeds under that contract. OpenAI recommends explicit required properties and additionalProperties: false for predictable records. The local pipeline should nevertheless validate downloaded content again, because transport handling, parsing mistakes, stale schemas or unexpected response states can still break the downstream contract.
Keep three outcomes distinct: request failure, unusable response structure and schema-valid classification. Only the third category proceeds to semantic scoring. A refusal, incomplete response, absent output text or non-success status is not equivalent to a schema-valid “unknown” classification unless the taxonomy explicitly defines such a label and the model actually returned it.
import json
from jsonschema import Draft202012Validator
schema = json.loads(
(run_dir / "classification-schema.json").read_text(encoding="utf-8")
)
validator = Draft202012Validator(schema)
def response_text(batch_row: dict) -> str:
response = batch_row.get("response") or {}
if response.get("status_code") != 200:
raise ValueError(
f"Non-success response status: {response.get('status_code')}"
)
body = response.get("body") or {}
text_parts = []
for item in body.get("output", []):
for content in item.get("content", []):
if content.get("type") == "output_text":
text_parts.append(content.get("text", ""))
if not text_parts:
raise ValueError("No output_text content found")
return "".join(text_parts)
def parse_classification(batch_row: dict) -> dict:
payload = json.loads(response_text(batch_row))
errors = sorted(validator.iter_errors(payload), key=lambda e: list(e.path))
if errors:
messages = [
f"{'/'.join(map(str, error.path)) or '<root>'}: {error.message}"
for error in errors
]
raise ValueError("; ".join(messages))
return payload
The exact response traversal should match the official software development kit (SDK)A collection of libraries, tools and documentation for building against a platform. Open glossary entry version and raw artefacts in use. Add fixture tests using sanitised downloaded examples rather than assuming a helper returns the same shape forever. Preserve the raw line alongside the parsed object so an SDK or parser upgrade can be audited.
Score format, labels and escalation behaviour separately
Start with coverage counts: total holdout records, successful request records, request errors, missing outcomes, parse failures and schema failures. A semantic score calculated only over successful records can conceal systematic failure on long, ambiguous or sensitive tickets. Report both the denominator for the entire holdout and the denominator for schema-valid classifications.
For each schema-valid result, compare the predicted primary category and review route against the human ground-truth fields. The evaluation metrics are exact category agreement, per-category precision and recall, macro-averaged results across categories, and a separate recall measure for records that humans marked as requiring review. These are evaluation calculations, not universal acceptance targets.
| Measure | Calculation | Operational question |
|---|---|---|
| End-to-end usable coverage | Schema-valid classifications ÷ all holdout records | How much of the holdout produced a contract-valid result? |
| Exact category agreement | Correct primary categories ÷ schema-valid classifications | How often did the model match the adjudicated category? |
| Category recall | Correct predictions for a category ÷ human-labelled records in that category | Which categories are being missed? |
| Category precision | Correct predictions for a category ÷ model predictions of that category | Which predicted queues attract unrelated tickets? |
| Review-route recall | Correctly escalated records ÷ human-labelled review-required records | How often does the workflow preserve human review where expected? |
| False safe-route count | Review-required records predicted as routine | How many sensitive cases would bypass the intended review boundary? |
Use the acceptance criteria set before results were available. Thresholds are risk decisions owned by the team responsible for the workflow, not defaults supplied by the model. An example gate can require zero unresolved format failures, minimum performance for every critical class, and manual approval of every false safe-route case. Do not relax a threshold after observing a disappointing run without documenting the reason and evaluating the revised rule on fresh data.
from collections import Counter, defaultdict
rows = [] # Populate from reconciled, schema-valid predictions and manifest labels.
total = len(rows)
exact = sum(r["expected_category"] == r["predicted_category"] for r in rows)
by_category = defaultdict(Counter)
for row in rows:
expected = row["expected_category"]
predicted = row["predicted_category"]
for category in {expected, predicted}:
if expected == category and predicted == category:
by_category[category]["tp"] += 1
elif expected != category and predicted == category:
by_category[category]["fp"] += 1
elif expected == category and predicted != category:
by_category[category]["fn"] += 1
false_safe_routes = [
row for row in rows
if row["expected_review"] is True
and row["predicted_review"] is False
]
report = {
"schema_valid_records": total,
"exact_category_agreement": (exact / total) if total else None,
"false_safe_route_count": len(false_safe_routes),
"categories": {},
}
for category, counts in sorted(by_category.items()):
tp, fp, fn = counts["tp"], counts["fp"], counts["fn"]
report["categories"][category] = {
"precision": tp / (tp + fp) if tp + fp else None,
"recall": tp / (tp + fn) if tp + fn else None,
"support": tp + fn,
}
Perform manual error analysis before changing the prompt
Quantitative scores identify where disagreement occurs; they do not explain why. Sample all false safe-route cases, all errors in critical categories, every request or schema failure, and a stratified selection of ordinary disagreements. Reviewers should see the authorised minimised source record, human label, predicted object, taxonomy rule, prompt version and raw Batch outcome.
Classify each disagreement into a controlled error taxonomy. Recommended categories include incorrect model classification, ambiguous taxonomy, inconsistent human label, insufficient source context, multiple valid issues, preprocessing loss, schema or parser defect, and review-policy mismatch. Do not automatically count every human–model disagreement as a model error. Ground-truth labels can be inconsistent, but any correction must be independently adjudicated rather than changed merely to match the prediction.
Adjudication rule: if two reviewers cannot resolve a disagreement from the authorised record and written taxonomy, mark the example ambiguous, route comparable production records to people, and revise the taxonomy before treating it as a stable training or evaluation example.
For each proposed fix, state the failure mechanism it addresses. A taxonomy revision may solve overlapping categories; a preprocessing change may restore omitted context; a prompt clarification may encode an existing tie-break rule; a human-review rule may be the correct response to irreducible ambiguity. Increasing reasoning effort is an experiment, not a default remedy, and OpenAI’s guidance calls for measuring its effect on representative examples.
Handle failed and expired work without losing provenance
OpenAI documents that Batch requests can fail or expire. Preserve the original request JSONL, manifest, Batch record, downloaded files and status history. Never rebuild an expired subset from a mutable live export: records may have changed, identifiers may no longer align, and the resubmission would cease to represent the evaluated dataset.
Create a resubmission set only from manifest records without a usable terminal result. Give the attempt its own run identifier and retain a parent reference to the original custom_id. A recommended pattern is to append a non-sensitive attempt suffix while preserving the stable source mapping in the manifest. Ensure that every custom_id in the new Batch is unique, including across attempts.
{
"custom_id": "ticket-hash-0042-attempt-2",
"parent_custom_id": "ticket-hash-0042-attempt-1",
"source_record_ref": "ticket-hash-0042",
"resubmission_reason": "original request expired",
"previous_batch_id": "<previous batch identifier>"
}
Do not resubmit successful records merely to make files line up, and do not merge multiple successful attempts by taking whichever label is preferred. Define a deterministic rule: the first usable result from the approved configuration is retained, later duplicates are quarantined, and conflicting usable attempts require human review. This prevents retry behaviour from becoming an undocumented form of majority voting or result selection.
Escalate uncertainty and sensitive classifications to people
Strict Structured Outputs validates shape, not truth, fairness, safety or business appropriateness. A schema-valid result can still apply the wrong category, cite weak evidence or route a sensitive ticket incorrectly. Human review therefore remains mandatory for legal issues, threats, account security, payment disputes, identity concerns, safety matters, vulnerable customers, policy exceptions and any other locally designated sensitive route.
Also escalate records with missing context, conflicting issues, an explicit review flag, unsupported evidence, parser anomalies, request errors, expired outcomes or taxonomy ambiguity. If the schema includes a model-produced confidence-like field, treat it only as descriptive model output unless the organisation has separately tested its calibration. It must not override deterministic escalation rules.
The reviewer interface should present the minimised source text, predicted labels, extracted evidence, applicable taxonomy rule, processing status and audit identifiers. It should allow an authorised person to accept, correct or mark the case unresolved while recording a reason. The reviewer, not the model, owns any consequential routing, reply, closure, priority, legal, security or customer-account action.
After corrections, store adjudicated outcomes separately from raw model output. Feed only reviewed, authorised examples into future holdout revisions, and prevent previously inspected tuning examples from silently entering the next supposedly unseen test set. This separation preserves the value of the holdout as an independent scale-up gate.
Make the scale-up decision explicit
A pass decision requires all predetermined conditions to be met: complete reconciliation or an approved explanation for every unresolved record, acceptable contract-valid coverage, category metrics above the chosen thresholds, satisfactory performance on critical classes, review-route behaviour within policy, and signed manual analysis of consequential errors. Record who approved the decision, which run they reviewed and what constraints remain.
A fail decision should stop full-export submission. Create a bounded change proposal, modify one identifiable part of the configuration where practical, and run a fresh evaluation. Repeatedly tuning against the same holdout risks tailoring the workflow to known examples, so reserve a final untouched set or refresh the holdout through independent human labelling before production approval.
Even after a pass, the output remains advisory. Submit the full export only as another offline Batch run, reconcile it by custom_id, process its output and error files, and apply the same human-review rules. Do not connect raw classifications directly to automated replies, ticket closure, priority changes or customer-account actions.
Operate the classifier as a controlled data pipeline
Once the holdout gate has passed, treat each full-export run as a controlled release rather than a routine model call. The operating record should connect the authorised source export, minimisation rules, taxonomy version, prompt version, JSON Schema version, model identifier or snapshot, reasoning configuration, Batch job identifier, output and error artefacts, reconciliation report, evaluation result, reviewer decisions, and final disposition. This lineage is necessary because OpenAI documents Batch processing as asynchronous and warns that output order does not necessarily match input order.
OpenAI’s GPT-5.5 model reference, accessed on 30 September 2026, lists support for the Responses API, Batch, and Structured Outputs. That capability statement does not establish that a particular classification is correct or suitable for operational use. Continue to separate three questions in monitoring: whether the request completed, whether the result satisfied the schema, and whether the classification met the human-defined quality and review rules.
Record one immutable run manifest
Operating control: write a run manifest before submission and append observations without silently replacing its original configuration fields. Store hashes rather than unnecessary duplicate ticket text where a hash is sufficient for integrity checking. Access to the manifest and associated artefacts should follow the same or stricter controls as the authorised source export.
{
"run_id": "run-YYYYMMDD-sequence",
"purpose": "offline_ticket_classification",
"authorisation_reference": "internal-approval-reference",
"source_export_hash": "sha256-placeholder",
"input_jsonl_hash": "sha256-placeholder",
"manifest_hash": "sha256-placeholder",
"taxonomy_version": "taxonomy-vN",
"prompt_version": "prompt-vN",
"schema_version": "schema-vN",
"model": "documented-model-or-pinned-snapshot",
"reasoning_configuration": "recorded-setting",
"holdout_version": "holdout-vN",
"holdout_gate_result": "pass|fail|not_run",
"batch_id": null,
"submitted_at": null,
"last_observed_status": "not_submitted",
"output_file_hash": null,
"error_file_hash": null,
"reconciliation_status": "pending",
"review_status": "pending",
"release_status": "blocked"
}
The example is a recommended local record, not an OpenAI response format. Never place an API key, access token, session cookie, customer password, or realistic secret in this file. If an operational ticket identifier is sensitive, preserve the previously established opaque mapping and restrict the reverse map to authorised staff.
Monitor completion, integrity, quality and review queues separately
A single “batch succeeded” indicator is inadequate. OpenAI documents a completion window of up to 24 hours for Batch, while also documenting failed and expired work. Completion therefore means only that the asynchronous processing stage reached a terminal state; it does not mean that every source record produced a usable, semantically correct classification.
| Monitoring layer | Evidence to collect | Decision rule | Human response |
|---|---|---|---|
| Submission | Input file identifier, Batch job identifier, endpoint, submission time and accepted request count | Block the run if identifiers or expected counts are absent | Verify whether the service accepted the file before attempting another submission |
| Lifecycle | Every observed status and timestamp, including a terminal failed or expired state | Do not describe a non-terminal job as complete | Continue conservative polling or invoke the failure runbook at a terminal state |
| Reconciliation | Expected, successful, errored, duplicate, unknown and missing custom_id counts |
The counts must reconcile and every expected identifier must have a recorded disposition | Investigate collisions, missing records and unexpected identifiers before release |
| Contract validity | JSON parsing, JSON Schema validation, allowed taxonomy values and required-field checks | Any invalid record is quarantined rather than coerced silently | Inspect the raw response and request; correct code or configuration before resubmission |
| Semantic quality | Holdout metrics, per-class errors, sensitive-route misses and reviewed production samples | Use the predeclared acceptance criteria; do not lower them after seeing a poor result without documented approval | Stop scale-up, analyse errors and re-evaluate a versioned change |
| Human review | Queue age, unresolved sensitive cases, reviewer overrides and reasons | No downstream use while required reviews remain unresolved | Allocate authorised reviewers or pause intake rather than bypassing the gate |
Reconciliation metric: calculate reconciliation coverage as the number of expected custom_id values with exactly one recorded disposition divided by the number of expected identifiers. A disposition may be a validated output, a recorded request error, or a documented terminal absence awaiting resubmission. Do not count an unknown identifier as coverage, and do not treat duplicate outputs as two successful source records.
Monitor reviewer overrides by taxonomy class and reason, but do not interpret a low override rate as proof of correctness. Reviewers may miss errors, sampling may be unrepresentative, and a schema-valid result can still be factually or operationally wrong. Periodically insert known holdout records or conduct blinded sampling according to an approved quality plan.
Control model, prompt, schema and taxonomy changes together
A reproducible classifier depends on more than the model name. The effective configuration includes the model or model snapshot, prompt, examples, JSON Schema, taxonomy, tie-break rules, reasoning effort, preprocessing, excluded-field rules, and result parser. A change to any one component can alter results and should create a new configuration version.
OpenAI’s GPT-5.5 guidance recommends evaluation on representative examples and measured tuning rather than assuming that a configuration change is beneficial. Where the model reference provides an appropriate snapshot, pin and record it for reproducibility. If a snapshot is changed, removed, or no longer appropriate, treat the replacement as a candidate release and rerun the independently maintained holdout before processing another full export.
Use a change matrix before deployment
| Proposed change | Minimum validation | Required approval |
|---|---|---|
| Prompt wording with unchanged contract | Schema checks, complete holdout, per-class comparison and sensitive-case review | Classifier owner and quality owner |
| Model identifier, snapshot or reasoning effort | Fresh holdout run using identical inputs, plus review of errors, completion behaviour and resource observations | Technical owner and risk owner |
| Taxonomy label or tie-break rule | Relabel affected holdout examples, adjudicate disagreements and test migration of downstream consumers | Business taxonomy owner |
| JSON Schema or parser | Positive and negative fixtures, strict validation, compatibility check and reconciliation dry run | Engineering owner |
| Source extraction or minimisation | Authorisation review, field-level privacy check and representative holdout regeneration where inputs changed | Data owner and privacy or security reviewer as applicable |
| Acceptance threshold or review rule | Documented risk rationale and retrospective application to the unchanged holdout | Accountable business and risk owners |
Never compare two configurations using differently labelled test sets without explaining the difference. Preserve the old holdout labels, record adjudicated corrections as a new holdout version, and report both the configuration change and label-set change. Otherwise, an apparent improvement may come from altered ground truth rather than model behaviour.
Do not depend on the retiring API Evals platform. Keep the holdout, scoring code, expected labels, adjudication notes and historical results in an independently maintained system. The pipeline may use OpenAI guidance on representative ground-truth evaluation without making the retiring platform an operational dependency.
Revalidate product lifecycle assumptions
GPT-5.5 is scheduled to retire from ChatGPT, ChatGPT Work and Codex on 14 October 2026, but not from the API. Before every planned configuration change or publication, check the current model reference, account access and lifecycle documentation rather than inferring API retirement from a product-surface notice.
Use a runbook that stops unsafe scale-up
Runbook: assign a named operator, reviewer lead and incident owner before submission. The operator may manage files and job status; the reviewer lead owns human adjudication; the incident owner decides whether to pause, rollback or retire a configuration. Separation reduces the risk that the person who made a change also waives its failed evaluation.
- Confirm authority and scope. Match the export to its approval, verify the intended purpose, and reject records outside the authorised date range, queue or data class.
- Verify minimisation. Confirm that excluded attachments, credentials, payment data, unnecessary personal data and unsupported record types did not enter the request file.
- Lock configuration. Record the model or snapshot, reasoning setting, prompt, taxonomy, schema, parser and code revision.
- Confirm the gate. Require a passing result from the same configuration on the current human-labelled holdout. A previous model or prompt result is not transferable.
- Validate artefacts. Recompute hashes, count JSONL, verify unique
custom_idvalues and compare the request set with the local manifest. - Submit once. Record the accepted file and Batch identifiers before any retry. If the client loses its local response, investigate server-side state instead of immediately creating a duplicate job.
- Observe asynchronously. Poll conservatively, recording state changes. Do not promise immediate completion or schedule an irreversible downstream action for exactly 24 hours after submission.
- Retrieve all artefacts. Preserve output and error files, raw metadata and hashes under authorised access controls.
- Reconcile identifiers. Join only by
custom_id. Quarantine duplicates, unknown identifiers, malformed lines and records with no terminal disposition. - Validate results again. Parse and validate the structured object locally even though strict Structured Outputs constrained the requested shape.
- Apply review routing. Route uncertainty, sensitive topics, missing context, escalation candidates and policy-defined samples to authorised people.
- Approve or block release. Record who approved the reviewed classification dataset and what limited downstream use is permitted. Model output alone must not send replies, close tickets, change priority or alter customer accounts.
- Archive evidence. Retain only what policy requires, with access restrictions and a documented disposal date.
For human-in-the-loop Codex questions and bounded approvals, see human-in-the-loop Codex approval controls.
Rollback without erasing evidence
Rollback means stopping use of a candidate configuration and restoring a previously approved configuration for future work. It does not mean deleting evidence of a failed run, relabelling candidate outputs as approved, or assuming that previously processed records can be safely replayed without checking downstream state.
Define rollback triggers in advance
- The current holdout misses a predeclared overall, per-class or sensitive-route acceptance criterion.
- Reconciliation reveals unknown, duplicate or missing identifiers that cannot be explained and contained.
- The parser accepts records that violate the intended JSON Schema or taxonomy.
- A privacy or authorisation review finds that prohibited data entered the request file.
- Reviewer overrides expose a systematic taxonomy, prompt or preprocessing error.
- The model, snapshot, prompt, schema or reasoning configuration differs from the approved manifest.
- A downstream system used classifications beyond the approved read-only or review-only boundary.
Rollback procedure: pause new submissions; mark the affected run as blocked; preserve input, output and error evidence; identify the first affected configuration version; prevent unreviewed outputs from reaching consumers; restore the last approved version only after confirming its current availability and permissions; rerun the current holdout; and require fresh approval before resuming. If source data or taxonomy changed, a former configuration may no longer be an acceptable fallback.
Do not automatically resubmit an entire failed Batch. First derive the unresolved custom_id set from the manifest, output and error files. Create a new request file only for records that are authorised, still needed and safe to retry. Give the resubmission a new run and Batch identifier while retaining a parent-run reference. This prevents duplicate classifications from being mistaken for additional source records.
Protect ticket data throughout the lifecycle
Use only an authorised, read-only export and minimise it before upload. Data minimisation should be field-specific: remove values that the classifier does not need, redact secrets and authentication material, exclude unsupported attachments, and avoid including a full conversation when a bounded authorised excerpt is sufficient. Pseudonymous identifiers reduce direct exposure but do not make ticket text anonymous.
Keep access decisions human-led and organisation-specific. Confirm the applicable OpenAI account arrangement, workspace or project controls, retention requirements, region requirements, contractual terms, and internal security review before processing personal, confidential, regulated or security-sensitive material. The supplied sources establish API capabilities; they do not provide a general compliance determination for a reader’s deployment.
- Separate the source export, API request file, opaque-identifier reverse map, output files and reviewer notes where practical.
- Grant each operator only the access needed for their role; reviewers do not necessarily need source-system credentials.
- Encrypt stored artefacts using approved organisational controls and avoid copying them into logs, issue trackers or chat channels.
- Sanitise error logging so that diagnostic messages do not reproduce complete ticket bodies.
- Set retention and deletion dates for local inputs, downloaded outputs, error files, manifests and backups.
- Record lawful and policy-approved handling instructions for data-subject requests, legal holds or incident response where applicable.
- Require manual handling for credentials, account takeover indicators, legal threats, health information, financial details and other locally defined sensitive classes.
Strict Structured Outputs constrains format, not truth, safety, fairness, confidentiality or business appropriateness. A schema-valid label remains a model-generated proposal until the pipeline’s evaluation and review controls permit its limited use.
Failure catalogue and operator response
| Observed failure | Likely control gap | Immediate action | Recovery condition |
|---|---|---|---|
| Batch remains non-terminal | Asynchronous completion was treated as immediate | Continue recorded polling; do not submit a duplicate merely because processing is still underway | Terminal state observed or documented incident decision made |
| Batch expires | Requests did not complete within the documented window | Preserve all artefacts and identify records without a disposition | Only the unresolved authorised set is placed in a separately tracked resubmission |
| Batch or individual request fails | Invalid request, service error or unsupported configuration may be present | Inspect the documented error information rather than guessing or stripping controls | Root cause is corrected and the changed configuration passes validation and, where relevant, the holdout |
| Output count differs from input count | Errors, expiry, malformed lines or mistaken positional matching | Set-reconcile expected identifiers against successful and errored identifiers | Every expected identifier has one understood disposition |
Duplicate custom_id |
Input-generation defect or merged artefacts | Block release and trace duplicates to the original manifest | Uniqueness is restored in a new validated input file |
Unknown custom_id in output |
Wrong output file, mixed run or corrupted lineage | Quarantine the record and verify file and run identifiers | Artefact provenance is established; otherwise reject the affected artefact |
| Structured object fails local validation | Extraction, parser, schema-version or response-handling defect | Retain the raw line and compare it with the exact submitted request and schema | Parser and fixtures pass; affected records are reprocessed under a versioned run if needed |
| Schema-valid but wrong category | Semantic error, ambiguous taxonomy, missing context or weak example coverage | Send to human adjudication and classify the error type | A versioned correction passes the unchanged or properly revised holdout |
| Sensitive ticket receives an ordinary route | Escalation rule or review routing failed | Stop operational use and review the affected population, not only the observed case | Risk owner approves corrected routing after targeted and full holdout checks |
| Downstream automation acts on raw output | Review-only boundary was not enforced technically | Disable the integration, preserve audit evidence and assess affected records | Human approval and least-privilege enforcement are restored and tested |
Troubleshoot from evidence, not prompt guesses
The output appears to be missing records
Compare identifier sets rather than line counts alone. Build sets for expected, successful, errored and unknown custom_id values. Then calculate missing identifiers as expected minus the union of successful and errored identifiers. Check the terminal Batch state and error artefact before deciding whether any request should be resubmitted. Output order is not evidence of source order.
Every line parses, but holdout quality fell
Confirm that the evaluated configuration exactly matches the candidate manifest. Compare model or snapshot, reasoning effort, prompt bytes, examples, taxonomy, schema, preprocessing and parser revision. Review errors by category and risk severity. Do not immediately add instructions for individual mistakes; first determine whether ground-truth labels conflict, context was removed, or the taxonomy lacks a deterministic rule.
The model returns allowed labels but reviewers disagree
Sample disagreements for blinded adjudication by at least the locally required number of qualified reviewers. Distinguish model error from ambiguous policy and inconsistent human labelling. If reviewers cannot apply the taxonomy consistently, revise the taxonomy and relabel affected holdout examples before tuning the model. A stricter schema cannot resolve a semantic policy dispute.
A retry may have created two jobs
Search the run ledger for the input hash, submission time, file identifier and any returned Batch identifier. Do not merge outputs until provenance is established. If both jobs processed the same identifiers, retain both artefacts for audit, choose one run according to a documented rule, and mark the other as duplicate processing. Never combine duplicate results by positional order or majority vote without an approved evaluation method.
The error rate changes after a parser update
Run the old and new parsers against a frozen fixture set containing valid outputs, malformed JSON, wrong types, unexpected properties, null cases, duplicate identifiers and representative errors. If only the parser changed, do not attribute the difference to GPT-5.5. Version the parser independently and regenerate reconciliation reports for affected runs where policy permits.
Frequently asked operational questions
Can a schema-valid result bypass human review?
No. OpenAI’s Structured Outputs guidance supports enforcing the requested JSON Schema shape, including required fields and rejection of unexpected properties through the documented schema pattern. It does not guarantee factual correctness, fair treatment, calibrated confidence or appropriate business action. Preserve review for sensitive, uncertain, escalation and consequential routes.
Should the full export be submitted immediately after a prompt change?
No. Freeze the candidate configuration, run the current representative holdout, inspect aggregate and per-class results, review sensitive failures, and obtain the predeclared approval. Submit the full export only if the candidate meets the acceptance criteria selected as an organisational risk decision.
Can completion be scheduled for exactly 24 hours?
No. OpenAI documents a 24-hour completion window for Batch, not a promise that every job finishes at exactly that time or succeeds. Design for asynchronous status monitoring, failed or expired requests, output and error retrieval, reconciliation, and controlled resubmission.
Is the OpenAI Evals platform required?
No. For this workflow, retain an independent holdout and local evaluation harness. OpenAI documents the Evals platform as deprecated, read-only from 31 October 2026 and shutting down on 30 November 2026. The holdout must remain usable after those dates.
Does the ChatGPT and Codex retirement notice end GPT-5.5 API use?
Not according to the documentation reviewed on 30 September 2026. OpenAI’s notice schedules GPT-5.5 retirement from ChatGPT, ChatGPT Work and Codex on 14 October 2026 and explicitly excludes the API. Recheck current API documentation and account access before each release because lifecycle and availability information can change.
Should failed records be resubmitted unchanged?
Only after identifying the failure cause. A transient processing failure may justify a controlled retry, while an invalid request, unsupported schema or privacy violation requires correction or exclusion. A modified request is a new versioned processing event and may require renewed holdout validation if the effective classifier changed.
Can output trigger ticket closure, replies or account changes?
Not in this tutorial’s operating boundary. Keep results in a reviewable staging dataset. Authorised people must decide any reply, closure, priority adjustment, legal escalation, security intervention or customer-account action under the organisation’s existing controls.
End-to-end release and operations checklist
Governance and data preparation
- Confirm the export is authorised, read-only and limited to the approved purpose.
- Record the accountable data owner, classifier owner, reviewer lead and incident owner.
- Remove unnecessary personal, confidential and security-sensitive fields.
- Exclude credentials, authentication material and unsupported attachments.
- Preserve an access-controlled mapping between source records and opaque identifiers only where required.
- Document retention, deletion, incident and legal-hold requirements.
Configuration and evaluation
- Version the taxonomy, tie-break rules, prompt, examples, JSON Schema, parser and preprocessing code.
- Record the documented GPT-5.5 model identifier or chosen snapshot and reasoning configuration.
- Maintain the human-labelled holdout outside the retiring API Evals platform.
- Include ordinary, ambiguous, multi-issue, missing-context and escalation-sensitive cases.
- Set acceptance criteria before inspecting candidate results.
- Run the exact candidate configuration against the complete current holdout.
- Review overall, per-class and sensitive-route errors rather than relying on one aggregate score.
- Obtain the required human approval before scale-up.
Input and submission
- Generate one unique
custom_idfor every request. - Validate every JSONL line and the embedded Responses API body locally.
- Confirm strict JSON Schema requirements and reject unexpected properties.
- Check that the request and manifest identifier sets match exactly.
- Hash the source, request file and manifest under approved controls.
- Load API credentials from an approved secret mechanism, never from source files or artefacts.
- Record the uploaded file and Batch job identifiers before retrying any client operation.
- Confirm that the Batch endpoint is
/v1/responsesfor this implementation.
Monitoring and retrieval
- Record submission time and each observed lifecycle state.
- Allow for asynchronous processing within the documented completion window.
- Handle failed and expired states without assuming all requests completed.
- Download and preserve both output and error artefacts when available.
- Record artefact hashes and maintain access restrictions.
- Avoid logging raw ticket bodies during polling or diagnostics.
Reconciliation and validation
- Join results to source records only by
custom_id. - Detect duplicate, missing and unknown identifiers.
- Give every expected identifier exactly one understood disposition.
- Parse and revalidate every structured result locally.
- Verify taxonomy membership, nullable fields and local business constraints.
- Quarantine malformed, failed, ambiguous and sensitive records.
- Do not substitute a default label merely to make counts balance.
Review and limited release
- Send policy-defined cases and samples to qualified human reviewers.
- Require reviewers to record overrides and reasons.
- Escalate legal, account-security, financial, health and other locally sensitive content.
- Investigate clusters of overrides as possible systemic errors.
- Keep all downstream use blocked until required reviews are complete.
- Permit only the approved read-only, analytical or staging use.
- Prevent model output alone from replying, closing, reprioritising or changing customer accounts.
Change, rollback and closure
- Rerun the holdout after any model, snapshot, reasoning, prompt, taxonomy, schema, parser or preprocessing change.
- Revalidate GPT-5.5 API lifecycle and account access against current documentation.
- Pause submissions when a rollback trigger occurs.
- Preserve failed-run evidence and identify the complete affected record population.
- Resubmit only unresolved, authorised records under a new linked run.
- Record final counts for successful, errored, missing, duplicated, reviewed and excluded records.
- Document final approval, permitted use, unresolved issues and disposal dates.
- Delete artefacts when the approved retention period ends, subject to applicable holds and policy.
Operational boundary: The Batch API is not suitable for real-time interaction because a batch can use a 24-hour completion window. A valid JSON Schema and Structured Outputs do not guarantee semantic correctness: every custom_id must be reconciled, the error file must be inspected, holdout results must be reviewed and uncertain or consequential classifications require human review.
Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!
Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.
Useful Links
- OpenAI: GPT-5.5 model reference
- OpenAI: Batch API guide
- OpenAI: Structured Outputs guide
- OpenAI: Using GPT-5.5
- OpenAI: Working with evals
- OpenAI: ChatGPT Work and Codex model lifecycle
