OpenAI Introduces GPT-6.1 Sol for ChatGPT Work, Codex, and the API: Availability, Pricing, and Limits

OpenAI Introduces GPT-6.1 Sol for ChatGPT Work, Codex, and the API: Availability, Pricing, and Limits

29 September 2026: OpenAI has introduced generative pre-trained transformer (GPT)A family of transformer-based language models trained to generate and analyze content. Open glossary entry-6.1 Sol, a new model identifier positioned as an upgrade to GPT-6 Sol for agentic coding, computer use and professional work. OpenAI’s announcement places the model in ChatGPT Work, Codex and the application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry, while explicitly excluding it from regular ChatGPT conversations. The distinction matters because seeing GPT-6.1 Sol documented for “ChatGPT” does not mean that it is available in an ordinary ChatGPT chat or to every paid account.

OpenAI’s release notes describe a staged rollout beginning with Pro and expanding to eligible paid plans. Its announcement, meanwhile, lists paid-plan availability in ChatGPT Work and Codex from 29 September 2026. Organisations and individual subscribers should therefore verify the model selector, workspace permissions and current rollout status in their own account rather than treating the announcement date as proof of immediate access.

The release also changes the API identifier to gpt-6.1-sol. OpenAI’s developer documentation lists standard text-token rates of $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. These are metered token rates, not a fixed price for completing a task, and they do not establish that GPT-6.1 Sol will be cheaper for a particular production workload.

GPT-6.1 Sol follows the earlier GPT-6 Sol release, but it should be treated as a separately dated model version rather than as an interchangeable label. Teams that pin model identifiers, run approval tests or maintain procurement inventories should record gpt-6.1-sol as its own deployment dependency and re-evaluate it under the conditions that matter to their applications.

Abstract amber sphere moving through a structured blue technical landscape
Abstract amber sphere moving through a structured blue technical landscape.

Evidence checkpoints

Documented point: OpenAI announced GPT-6.1 Sol on 29 September 2026 as an upgrade to GPT-6 Sol for agentic coding, computer use and professional work. This is OpenAI product positioning, not an independently established outcome for every workflow. [official source 1 official source 2]

Documented point: GPT-6.1 Sol is for ChatGPT Work and Codex rather than regular ChatGPT conversations; access depends on plan, workspace settings, and rollout status. Release notes describe rollout beginning with Pro; Enterprise and Education access can be governed by workspace permissions. [official source 1 official source 2]

Documented point: The documented API model identifier (ID)A value used to distinguish one record, task, source or object from another. Open glossary entry is gpt-6.1-sol, with standard rates of $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. Token rates are not a fixed task price; cache writes, long context, Fast mode, Batch/Flex, regional processing, reasoning, tools, and commercial arrangement can affect charges. [official source 1]

Documented point: Documentation lists text and image input, text output and a 1,050,000-token context window, does not list audio or video modalities, and recommends the Responses API for tool calling. Capacity and feature support do not guarantee task success or production suitability. [official source 1]

Documented point: OpenAI reports favourable benchmark and alignment results versus GPT-6 Sol and selected comparators, while its announcement and system card limit those results to evaluated settings. Evaluated environments may differ from production ChatGPT; several tests use difficult or failure-eliciting prompts and are not typical-use evidence. [official source 1 official source 2]

GPT-6.1 Sol: verified release facts at a glance

The following table consolidates the product announcement dated 29 September 2026, OpenAI’s release notes from the same date, the Help Center article updated on 30 September and the developer model reference accessed on 30 September. Availability remains account- and workspace-dependent, so the table describes OpenAI’s documented product boundaries rather than guaranteeing access in any particular account.

Question Documented position Operational qualification
What was announced? OpenAI announced GPT-6.1 Sol on 29 September 2026 as an upgrade to GPT-6 Sol, positioned for agentic coding, computer use and professional work. This is OpenAI’s product positioning. It does not independently establish better results for every task, organisation or deployment.
Where is it documented as available? OpenAI identifies ChatGPT Work, Codex and the API as supported product surfaces. These surfaces have different access paths, controls and billing arrangements. Availability on one does not imply availability on another.
Is it available in regular ChatGPT conversations? No. OpenAI’s cited documentation explicitly separates GPT-6.1 Sol from regular ChatGPT conversations. Users should not expect the model merely because they can open an ordinary ChatGPT chat or hold a paid ChatGPT subscription.
How is the rollout described? OpenAI’s announcement lists paid-plan availability in ChatGPT Work and Codex from 29 September, while its release notes describe rollout beginning with Pro and expanding to eligible paid plans. Check the actual account, workspace and model selector. A staged rollout can produce different access states among otherwise similar users.
What is the API model identifier? gpt-6.1-sol Do not silently substitute the earlier GPT-6 Sol identifier. Record the selected model version in evaluation and deployment evidence.
What are the listed standard token rates? $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. Token mix, cache behaviour, reasoning effort, tools, long context, processing mode, region and commercial terms can change the effective cost.
What inputs and outputs are documented? OpenAI lists text and image input with text output. The model reference does not list audio or video as supported modalities. Listed support does not guarantee successful interpretation of every image or document.
What context capacity is listed? A 1,050,000-token context window. Capacity is not evidence that the model will use every part of a long context accurately or that maximum-context requests are economical.
Which API approach does OpenAI recommend for tool use? OpenAI recommends the Responses API for tool calling. Tool permissions, input validation, output checks and human approval remain application responsibilities.
What performance evidence is available? OpenAI reports favourable results against GPT-6 Sol and selected comparators in its announcement and deployment-safety material. These are vendor-published evaluations under defined conditions, not independently replicated proof of production quality, factuality or safety.

Why the product-surface boundary matters

“ChatGPT Work”, “Codex” and “regular ChatGPT conversations” are not interchangeable names for the same interface. OpenAI’s Help Center documentation separates ChatGPT Work and Codex from ordinary ChatGPT chat, and specifically states that GPT-6.1 Sol is unavailable in regular conversations. A team assessing access must therefore identify the exact surface on which the work will run before comparing capability, cost or governance.

ChatGPT Work is the documented professional-work surface associated with this release. According to OpenAI, access depends on the user’s plan, rollout state and, where applicable, workspace configuration. Enterprise and Education workspaces can also place model access behind permissions controlled by administrators. A model’s appearance in public documentation does not override a disabled workspace setting or establish entitlement for every member.

Codex is the coding-oriented surface named in the announcement and release notes. Its inclusion supports OpenAI’s positioning of GPT-6.1 Sol for agentic coding, but it does not establish that every Codex client, local configuration, cloud environment or account receives the model simultaneously. Teams should confirm the model actually offered in their Codex environment and preserve the selected identifier in task or evaluation records where reproducibility matters.

Regular ChatGPT conversations are expressly outside the documented availability boundary. A user opening a standard conversation should not infer that an automatic or unnamed backend selection is GPT-6.1 Sol. Nor should an organisation describe the model as broadly available “in ChatGPT” without adding the Work qualification, because that wording erases a restriction made explicit by OpenAI.

The API is a separate developer surface with the model identifier gpt-6.1-sol and metered token pricing. API availability does not establish ChatGPT Work or Codex access, just as a visible Work model does not establish that an API project has permission to call it. Developers should check project access and current documentation before changing production configuration.

A practical access check

Recommended verification procedure: first name the intended surface as ChatGPT Work, Codex or API. Next, confirm the account or organisation is on an eligible plan without assuming that “paid” means already enabled. For a managed workspace, ask an authorised administrator to verify model permissions and role-based access control (RBAC)A method of assigning permissions through defined roles rather than one person at a time. Open glossary entry. Finally, inspect actual availability in the relevant product and record the date because OpenAI describes the release as staged and documentation can change.

  • For ChatGPT Work: confirm that the user is operating inside the intended workspace, not a personal or regular ChatGPT conversation, and check whether GPT-6.1 Sol appears as an available model.
  • For Codex: verify the current client or environment, authentication route and offered model list rather than relying on another user’s screenshot or account state.
  • For the API: confirm that the project can select gpt-6.1-sol, then test with non-sensitive representative requests before changing production traffic.
  • For Enterprise or Education: confirm workspace permissions with the responsible administrator because OpenAI says access can be governed by workspace controls.

A failed access check does not by itself show that the announcement is inaccurate. It may reflect the staged rollout, an ineligible plan, a workspace restriction, a different product surface or an account-specific access condition. Conversely, one user seeing the model does not establish organisation-wide availability.

Published API pricing is a rate card, not a completed-task price

OpenAI’s model reference lists standard rates of $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. The listed cached-input rate may be material to applications that repeatedly send eligible shared prefixes, but it should not be applied to all input by default. Only tokens treated as cached input under the applicable service conditions receive that listed rate.

A basic token-only estimate can be written as follows. This is a calculation template, not a quotation:

estimated token charge =
  (uncached input tokens / 1,000,000 × $2)
+ (cached input tokens / 1,000,000 × $0.10)
+ (output tokens / 1,000,000 × $10)

For example, an internal estimate must separate uncached and cached input rather than multiplying all prompt tokens by the cheapest listed rate. It must also count model-generated reasoning and output according to the applicable API reporting and billing rules. OpenAI’s published unit prices do not promise a specific token mix, cache-hit rate, response length or number of tool calls.

Several variables can make a task’s effective cost differ from a simple rate-card estimate. OpenAI’s documentation identifies conditions involving cache writes, long context, Fast mode, Batch processing or Flex processing, regional processing and commercial arrangements. Reasoning effort, tool execution, retries and application design can also affect total usage. Readers should consult the live documentation and their applicable commercial terms before approving a budget.

The appropriate comparison is therefore cost per accepted workflow result, not input-token price alone. That measure should include failed or incomplete requests, retries, tool calls, output length, human review and any additional infrastructure needed to complete the task. None of the sources cited here establishes that GPT-6.1 Sol costs one-fifth as much for every workload, lowers completed-task cost or is automatically the least expensive Sol-family option.

Decision rule: use the published rates to build a workload-specific estimate, then validate that estimate with representative, authorised traffic. Do not convert a cached-input discount or a per-million-token rate into a universal claim about per-task savings.

Documented capabilities and their limits

OpenAI’s developer reference lists text and image input, text output and a 1,050,000-token context window. It does not list audio or video input or output for GPT-6.1 Sol. The model documentation also recommends the Responses API when an application needs tool calling.

The context figure is a capacity specification, not a measure of reliable recall, reasoning quality or factual consistency across one million tokens. Long inputs can contain contradictory instructions, irrelevant records, stale material or adversarial content. Production applications should test retrieval and citation behaviour at realistic document lengths instead of assuming that material is understood merely because it fits within the context window.

Image-input support likewise identifies an accepted modality rather than guaranteeing accurate extraction from every diagram, screenshot, scan or visually complex document. Where image interpretation affects a financial, legal, health, employment, security or other consequential action, a qualified person should inspect the source and verify the model’s output before action is taken.

Tool-calling support should be treated as an integration capability, not delegated authority. Applications remain responsible for authenticating users, limiting tool permissions, validating arguments, constraining destinations, logging material actions and placing human approval in front of consequential writes. A model capable of computer use can still misunderstand page state, select an incorrect control or follow untrusted instructions embedded in content.

Recommended adoption test: create a fixed set of representative tasks for the intended Work, Codex or API workflow. Record input versions, model identifier, reasoning configuration, enabled tools, permissions, expected result and reviewer decision. Test ordinary cases alongside ambiguity, missing context, tool failure, long-context conflict and attempts to redirect the model. Retain GPT-6.1 Sol only if the measured configuration clears the organisation’s pre-set acceptance criteria.

How to read OpenAI’s benchmark and safety claims

OpenAI reports favourable benchmark, factuality and alignment results for GPT-6.1 Sol relative to GPT-6 Sol and selected comparators. Those results are vendor-published evaluations. They should be attributed to OpenAI and read with the task definitions, prompts, model settings, tools and scoring conditions supplied for each evaluation.

The announcement and deployment-safety addendum do not convert those evaluations into a general guarantee of real-world accuracy or safety. OpenAI notes that evaluated environments can differ from production ChatGPT, and some safety tests deliberately use difficult or failure-eliciting prompts. Such tests can be useful for examining a particular risk boundary, but they are not evidence of typical user behaviour or proof that harmful outcomes cannot occur.

Results reported for DeepSWE, GDP.pdf, AutomationBench, OSWorld, Terminal-Bench Science or factuality and safety suites should not be presented as independently replicated unless an independent replication is actually available. They also should not be combined into a single claim that GPT-6.1 Sol is universally superior to GPT-6 Astra or any other model. Different evaluations measure different tasks under different conditions.

OpenAI’s system-card risk categorisations and documented safeguards provide deployment context, not an assurance of compliance or an elimination of misuse risk. Organisations remain responsible for their own threat modelling, data handling, access controls, monitoring, incident response and sector-specific review. A system-card classification cannot determine whether a particular deployment satisfies an organisation’s contractual, regulatory or legal obligations.

What professionals should record before adoption

Because the release introduces a distinct model identifier and a staged access path, a deployment record should capture more than the name “Sol”. The minimum useful record is the exact surface, account or workspace, model identifier, access-check date, configuration, tools, data classification, token assumptions, evaluation results and approving owner.

  • Surface: ChatGPT Work, Codex or API; never the ambiguous label “ChatGPT” on its own.
  • Access evidence: the date access was checked, the eligible account or workspace and any administrator-controlled permission.
  • Model identity: gpt-6.1-sol for API work, plus any version or configuration information exposed by the relevant product.
  • Workload boundary: the task types, data classes, tools and actions permitted or prohibited.
  • Cost basis: uncached input, cached input, output, reasoning, tool, retry, context and processing assumptions.
  • Evaluation basis: representative test cases, acceptance criteria, known failures and the date of review.
  • Human authority: the person or role responsible for approving consequential outputs and production changes.

As of 30 September 2026, the defensible summary is narrow: OpenAI has announced GPT-6.1 Sol for ChatGPT Work, Codex and the API; regular ChatGPT conversations are excluded; rollout is staged and access-dependent; the API uses gpt-6.1-sol; and the published specifications and token rates require workload-specific validation. OpenAI’s evaluation results inform that assessment, but they do not replace independent testing, operational controls or human accountability.

What the model specification and rate card actually cover

OpenAI’s developer documentation, accessed on 30 September 2026, identifies the API model as gpt-6.1-sol. That identifier matters operationally: it distinguishes the 29 September GPT-6.1 Sol release from the earlier GPT-6 Sol model and should be recorded explicitly in deployment configurations, evaluation results, usage exports and change-control records. A product label such as “Sol” is not sufficiently precise when several generations may remain visible across different OpenAI surfaces.

OpenAI positions GPT-6.1 Sol as an upgrade to GPT-6 Sol for agentic coding, computer use and professional work. That description is vendor positioning rather than evidence that it will improve every coding, browser, document or tool-using workflow. Teams considering the model should test the exact model identifier against representative tasks and retain the prompts, tool definitions, reasoning settings, model responses and human judgements used in that comparison.

Layered model access paths represented by geometric blue and orange forms
Layered model access paths represented by geometric blue and orange forms.

Inputs, outputs and context capacity

The model reference lists text and image input with text output. It does not list audio or video as supported input or output modalities. An application that receives speech, recorded meetings or video therefore needs a separate, documented conversion or processing stage rather than assuming that gpt-6.1-sol can consume those formats directly.

Specification Documented position Operational implication
Model identifier gpt-6.1-sol Store the exact identifier with test results, requests and deployment approvals.
Input modalities Text and images Validate image legibility, ordering and task-specific extraction separately.
Output modality Text Any audio, visual or executable artefact must be generated or rendered by another controlled component.
Unsupported modalities Audio and video are not listed Do not send those formats on the assumption that the model will interpret them.
Context window 1,050,000 tokens Treat this as a capacity limit, not proof that every fact in a very long request will be used correctly.

OpenAI documents a 1,050,000-token context window. This permits unusually large requests, but context capacity is not the same as dependable retrieval, complete document coverage or correct synthesis. A production evaluation should place important evidence at different positions in long inputs, include conflicting and irrelevant material, and check whether the model cites or uses the intended source passages. Teams should also record input-token counts because OpenAI’s documented long-context conditions may change the applicable rate.

Images require comparable scrutiny. Listed image-input support establishes that the modality can be supplied; it does not establish perfect optical character recognition, chart interpretation, spatial reasoning or extraction from scans. Where an image influences a consequential decision, a human should compare the response with the original image and the application should preserve a traceable reference to the source page or asset.

A large context window can also create a data-governance problem if developers respond by placing every available record into one request. A recommended workflow is to minimise the source set before transmission, apply the organisation’s authorisation and retention rules, and separate unrelated users or cases. Capacity should not be treated as permission to combine personal, confidential or regulated information.

Endpoints, tools and reasoning controls

OpenAI’s model reference documents endpoint and tool support and recommends the Responses API for tool calling. For a new tool-using integration, that recommendation is the appropriate starting point. It does not mean that an existing integration can change only the model string: request fields, tool schemas, stopping behaviour, response handling and error recovery still require validation against the current model reference.

OpenAI’s model reference does not enumerate every supported endpoint and tool combination in a form summarised here, so no complete compatibility matrix is given. Before implementation, an engineer should consult the current model reference, record the endpoint and each required tool as separate acceptance criteria, and test only combinations that the documentation currently marks as supported. This prevents a broad statement such as “tool support” from being mistaken for support for every tool, endpoint or legacy request pattern.

Recommended tool test: run at least one expected call, one malformed argument set, one unavailable dependency, one tool timeout and one task in which no tool should be called. Confirm that the surrounding application validates arguments, restricts permissions, handles duplicate or incomplete calls, and does not treat model-generated tool output as trusted. The model should not receive broader file, network, account or execution rights merely because the tool interface permits them.

The developer documentation also lists reasoning settings for GPT-6.1 Sol. The documentation cited here does not provide the individual setting names, defaults or availability conditions, so those details should be read from the live reference rather than reconstructed from another model’s controls. Teams should save the selected reasoning configuration with each evaluation run because changing it may alter output quality, latency, token consumption and tool behaviour.

Recommended reasoning comparison: select only currently documented settings, keep the input, tools and acceptance rubric fixed, and compare task completion, unsupported claims, output tokens, tool calls, retries, elapsed time and reviewer acceptance. Retain the lightest setting that meets the defined quality and risk threshold. Do not assume that a higher reasoning setting is universally better or that a setting visible in one OpenAI product is exposed identically through the API.

Workflows that involve computer use or agentic action need a stricter boundary than ordinary drafting. Model support for a tool does not authorise purchases, publication, deletion, account changes, security actions or other consequential operations. A recommended control is to separate planning from execution, constrain tools to the minimum required scope, preview material changes and require an authorised person to approve high-impact actions.

How the published token prices should be interpreted

OpenAI’s model reference lists standard prices of $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. These are token rates, not a subscription price, a fixed charge per task or a prediction of the total cost of completing a workflow.

Token category Published standard rate What must be measured
Input $2 per million tokens Uncached instructions, user content, tool definitions and other billable request material.
Cached input $0.10 per million tokens Only tokens that qualify as cached under the applicable caching rules.
Output $10 per million tokens Billable generated tokens, including the effects of the chosen workflow and reasoning configuration.

A basic estimation formula is useful, provided it is labelled as an estimate:

estimated standard token charge =
  (uncached input tokens / 1,000,000 × $2)
+ (eligible cached input tokens / 1,000,000 × $0.10)
+ (output tokens / 1,000,000 × $10)

This formula covers only the three published standard token rates. It should not be used as a final invoice forecast where cache writes, long-context pricing, Fast mode, Batch processing or Flex processing, regional processing, tools, retries, failed work or negotiated commercial terms apply. The live documentation and the organisation’s agreement remain controlling.

Cached-input pricing also does not mean that any repeated text automatically receives the cached rate. Cache eligibility depends on the documented caching behaviour and request construction. Engineers should inspect actual usage records rather than estimating savings solely from the number of repeated prompt characters.

Cache writes and repeated prefixes

The rate card distinguishes cached input from ordinary input, while the pricing conditions also require attention to cache writes. A cache write and a cache read are not interchangeable billing events. Before modelling a recurring workload, identify how often a stable prefix is written, how often it is reused successfully, how frequently it changes and which tokens remain uncached.

Recommended cache model: divide traffic into first-use, successful-reuse, changed-prefix and non-eligible requests. Use measured token counts for each group and apply the currently documented write and read treatment. A low cached-input rate can have little effect where prompts vary early, reuse is sparse or large dynamic sections dominate the request.

Long-context rates

The model’s 1,050,000-token context window should not be interpreted as one flat price for every request size. OpenAI’s pricing documentation contains long-context conditions, and the applicable threshold and rates must be checked on the live model page before a cost forecast is approved. This article does not state those values, so no threshold or multiplier is inferred here.

Long requests can affect more than the input line item. They may produce longer outputs, more reasoning, additional tool calls or retries when evidence is difficult to locate. A cost comparison should therefore include the complete accepted task, not merely the first model request.

Fast, Batch and Flex processing

OpenAI documents pricing conditions for Fast mode and for Batch and Flex processing. Those categories must not be collapsed into the standard rate table. The documentation cited here does not establish that every reader has access to each mode, nor does it state their current rates or service conditions. Verify eligibility, processing characteristics and live pricing before selecting one for production.

Fast mode should not be presented as generally available or as a guaranteed latency result. Likewise, Batch and Flex are processing arrangements rather than evidence that a workload will be cheaper after retries, waiting constraints and operational handling are included. A procurement estimate should identify the exact processing mode, expected token mix and service requirements instead of applying an undocumented discount.

Regional processing and commercial terms

The developer documentation includes regional-processing price qualifiers. A regional option may have different pricing or availability, but this section does not state a percentage or region list because the sources reviewed do not supply one. Organisations with residency, localisation or contractual requirements should have the responsible privacy, security, procurement and legal teams verify that the selected processing option meets the actual requirement and that the correct rate appears in the agreement.

Published public rates may also differ from an organisation’s commercial arrangement. Budget owners should reconcile estimates with the applicable contract, account configuration, taxes and usage records. None of the documented token rates establishes the price of a ChatGPT plan, a Codex entitlement or a ChatGPT Work seat.

A task-cost calculation that avoids false precision

Recommended method: sample a representative set of completed tasks and record uncached input, cached input, cache-write treatment, output, reasoning configuration, context band, processing mode, tool charges where applicable, retries and failures. Calculate the median and upper-range accepted-task cost from observed usage. Keep rejected or incomplete attempts in the total because they still affect the economics of the workflow.

A claim that GPT-6.1 Sol costs “one-fifth as much” is not supported for an arbitrary workload merely because the cached-input rate is one-twentieth of the ordinary input rate. The effective ratio depends on the proportion of tokens that qualify for caching, the output share, cache-write treatment, long context, reasoning, tools and retries. Cost claims should state their workload, measurement period and assumptions.

Access is staged and surface-specific

OpenAI’s 29 September 2026 announcement lists GPT-6.1 Sol for eligible paid-plan use in ChatGPT Work and Codex. The release notes from the same date describe a staged rollout beginning with Pro and expanding to eligible paid plans. These statements are not contradictory if the model is being enabled progressively, but neither supports a claim of immediate access for every paid account.

The Help Center article updated on 30 September separates ChatGPT Work and Codex from regular ChatGPT conversations. GPT-6.1 Sol is explicitly unavailable in regular ChatGPT conversations under the cited documentation. A user who cannot select it in an ordinary conversation should not infer an account fault, and an administrator should not promise that purchasing a paid plan will place it in that surface.

Access rule: confirm the product surface first, then the plan, rollout status and workspace permissions. Do not use general ChatGPT access as evidence of ChatGPT Work, Codex or API entitlement.

For ChatGPT Work, the relevant check is whether the model appears within the Work experience available to that account and workspace. For Codex, the check is whether the account, Codex surface and current task configuration expose the model. For the API, the check is whether the project can use gpt-6.1-sol through the required supported endpoint. Success in one of these checks does not prove access through the others.

Enterprise and Education workspaces may be governed by workspace permissions. Workspace role-based access control can therefore affect who can use relevant Work or Codex capabilities even after OpenAI has rolled them out. Administrators should inspect actual workspace controls and avoid relying on screenshots or instructions from a personal account.

Rollout status can change after publication. The dates in this section describe OpenAI documentation issued on 29 September and updated or accessed on 30 September 2026; they are not a permanent availability guarantee. Recheck the release notes, Help Center and model reference before procurement, migration or user communications.

A deployment verification checklist

The following is a recommended verification procedure rather than an OpenAI guarantee. Record each answer with a date, account or workspace scope, responsible owner and evidence reference. A blank answer should block production assumptions rather than be treated as an implicit “yes”.

  1. Identify the surface. State whether the proposed workflow uses ChatGPT Work, Codex or the API. If it requires regular ChatGPT conversations, the cited documentation does not support GPT-6.1 Sol for that use.
  2. Confirm live access. Check the actual account, project and workspace. Record whether access is present, absent or still rolling out; do not infer it from paid status alone.
  3. Review workspace governance. For Enterprise or Education, have an authorised administrator verify relevant workspace permissions and RBAC rather than asking users to bypass controls.
  4. Pin the model identity. For API work, record gpt-6.1-sol in configuration, logs and evaluation records. Do not substitute a friendly label that obscures the version.
  5. Validate modalities. Confirm that the workflow supplies only documented text or image input and expects text output. Add separately authorised components for unsupported media.
  6. Check the endpoint. Verify the current model reference for the intended endpoint. Use the Responses API as OpenAI recommends for tool calling, while testing the complete request and response contract.
  7. Enumerate tools. List every required tool individually and confirm current model support. Apply least privilege, argument validation, timeouts, audit records and human approval for consequential actions.
  8. Select reasoning deliberately. Read the current documented settings and defaults. Evaluate available configurations on representative tasks instead of copying settings from another model or product surface.
  9. Test context behaviour. Measure token counts and test important evidence throughout long inputs. Do not equate the 1,050,000-token window with complete recall or correct synthesis.
  10. Separate price categories. Record uncached input, cached input and output tokens, then identify cache-write, long-context, Fast, Batch, Flex, regional and contractual qualifiers.
  11. Measure accepted-task cost. Include retries, failures, tool use and human review. Avoid presenting the public token rates as a guaranteed completed-task price.
  12. Run representative evaluations. Include expected cases, difficult cases, tool failures, conflicting evidence and tasks that should be refused or escalated. Keep human-authored acceptance criteria.
  13. Review risk controls. Treat OpenAI’s system-card categorisations and safeguards as evidence about evaluated conditions, not proof that harmful use is impossible or that a deployment is compliant.
  14. Set monitoring and rollback. Define who reviews quality, cost, tool behaviour and incidents; specify the conditions for disabling the model or returning to the previous configuration.
  15. Recheck before launch. Revisit the official announcement, release notes, Help Center, model reference and safety addendum because rollout, pricing and documentation can change.

OpenAI reports favourable benchmark, factuality and alignment results against GPT-6 Sol and selected comparators, but those are vendor-published evaluations under defined conditions. Several cited evaluations use difficult or failure-eliciting prompts, and evaluated environments may differ from production ChatGPT, Codex or a reader’s API application. They can inform test design, but they do not replace workload-specific evaluation or establish universal accuracy, superiority or safety.

The deployment-safety addendum similarly provides risk categorisations, safeguards and evaluation context rather than an assurance that GPT-6.1 Sol cannot enable harm. A production decision should combine that material with the organisation’s own threat model, permissions, data classification, review procedures and incident controls. Compliance and legal suitability remain organisation-specific human decisions.

What OpenAI’s evaluations can—and cannot—establish

OpenAI’s 29 September 2026 announcement positions GPT-6.1 Sol as an upgrade to GPT-6 Sol for agentic coding, computer use and professional work. The announcement and the accompanying deployment-safety addendum report favourable benchmark and alignment results against GPT-6 Sol and selected comparators in evaluated settings. These are vendor-published findings, not independently replicated tests, and they should be treated as evidence about performance under OpenAI’s stated evaluation conditions rather than proof that the model will outperform alternatives in a particular organisation.

The distinction matters because an evaluation result depends on the task set, scoring method, prompt construction, tool environment, model configuration and failure-handling rules. A model can score well on a bounded coding task yet still produce an unsuitable patch in a large repository with undocumented conventions. It can complete a computer-use benchmark while failing when an organisation’s application introduces custom authentication, inaccessible controls, transient errors or ambiguous confirmation steps. It can also perform strongly on professional-work questions without reliably applying a company’s current policies or proprietary terminology.

OpenAI’s cited materials also distinguish evaluated environments from production product use. Results reported for a model in a controlled harness must not be assumed to describe the behaviour of ChatGPT Work, Codex or an API integration in every configuration. Product instructions, available tools, workspace permissions, context supplied by the user, reasoning effort and surrounding software can all change the effective system being tested.

Measured benchmark and rollout evidence represented by balanced abstract structures
Measured benchmark and rollout evidence represented by balanced abstract structures.

Coding evidence should be tested against the real repository

OpenAI reports coding evidence that includes DeepSWE and other software-engineering evaluations. Those results can support a decision to test GPT-6.1 Sol for code investigation, patch drafting or tool-assisted engineering, but they do not establish that generated changes are correct, secure or maintainable in an organisation’s own codebase. Repository scale, dependency state, test quality, build tooling, language mix and local design rules are material variables that a benchmark cannot fully represent.

A suitable coding evaluation should use authorised, representative tasks drawn from recent work rather than a collection of unusually clean examples. Teams should include straightforward fixes, incomplete bug reports, cross-file changes, dependency conflicts, stale documentation and cases where the correct action is to request clarification. If a workflow gives Codex access to tools, the test should use the same practical tool boundaries and approval controls planned for deployment; testing a tool-rich configuration and deploying a more restricted one would not support the same conclusion.

Recommended coding test: prepare a blinded set of previously resolved issues and preserve the accepted patches, tests and review notes as reference material. Ask GPT-6.1 Sol to diagnose each issue and propose a patch without exposing the accepted solution. Authorised engineers should then assess functional correctness, change scope, test coverage, security implications, adherence to repository conventions and whether the model identified uncertainty. The model-generated patch must remain a draft until a human reviewer has inspected it and the organisation’s normal checks have passed.

Pass criteria should reflect release risk rather than a single aggregate score. A team might require no unauthorised file changes, no invented dependencies, successful execution of relevant tests, and human acceptance of both the diagnosis and patch. That is an example framework, not a universal threshold. Security-sensitive, safety-critical or regulated software requires stricter controls and specialist review, even when the model performs well on ordinary maintenance tasks.

Failure analysis is more useful than counting accepted patches alone. Record whether failures arose from misunderstood requirements, missing context, incorrect tool use, excessive changes, weak tests, fabricated interfaces or failure to stop before a consequential action. A model that solves many routine tasks but occasionally changes access-control logic without clear justification may be inappropriate for autonomous operation despite a favourable average result.

Professional-work results require organisation-specific source packs

OpenAI reports professional-work evaluations including GDP.pdf. Such tests can indicate how the model handled the documents and scoring conditions used by OpenAI, but they do not prove accuracy across legal, financial, procurement, policy, engineering or management work. “Professional work” covers tasks with markedly different evidence requirements and consequences, so an overall result cannot substitute for a departmental evaluation.

Recommended professional-work test: select completed tasks whose authoritative sources and accepted outputs can be reviewed. Examples include drafting a project status summary from approved records, extracting requirements from a controlled specification, comparing policy versions, or preparing a non-final analysis for a qualified reviewer. Remove unnecessary personal or confidential information, retain document versions, and require the model to distinguish sourced statements from assumptions and missing evidence.

Reviewers should score factual support, coverage, interpretation, source traceability, handling of contradictions and compliance with the requested output contract. Fluency should not compensate for unsupported claims. A polished answer that silently resolves conflicting documents is a more serious failure than an answer that explicitly asks for clarification.

Any task that affects employment, credit, insurance, healthcare, legal rights, procurement awards, financial commitments or access to essential services must remain human-led. GPT-6.1 Sol may organise evidence or draft material where organisational policy permits, but a qualified and authorised person should verify the sources, apply governing rules and make the consequential decision. OpenAI’s evaluation claims do not establish legal compliance, professional competence or suitability for a specific regulated process.

Computer-use scores do not remove action-level risk

OpenAI reports results involving AutomationBench and OSWorld, which concern computer-use and task-execution capabilities. These findings are relevant to workflows in which a model navigates software, operates interfaces or coordinates tools. They are not evidence that it can safely control every application, recognise every state transition or recover correctly from every interruption.

Computer-use tests should cover the exact application versions, permissions and interface patterns intended for deployment. Include expired sessions, duplicate buttons, delayed page updates, inaccessible controls, unexpected modal windows, partial saves and ambiguous success messages. The evaluation should also verify whether the model stops when the interface differs from its instructions rather than improvising a potentially consequential action.

Recommended safeguard: classify actions before testing. Read-only navigation and draft preparation can occupy a lower-risk class. External communications, purchases, account changes, data deletion, permission changes, software deployment and irreversible submissions belong in higher-risk classes. Require explicit human confirmation immediately before any high-risk action, and ensure the confirmation identifies the target, proposed change and material consequences rather than presenting a generic approval button.

Use accounts and environments specifically prepared for evaluation. Limit permissions, use non-production data where feasible, block unnecessary destinations and make rollback possible. A successful completion rate does not reveal whether the model attempted actions outside scope, exposed sensitive data in a tool call or reached the correct outcome through an unsafe path. Logs should therefore capture attempted actions, tool results, approvals, errors and final state, subject to applicable privacy and retention rules.

Science evaluations indicate a test candidate, not scientific authority

OpenAI includes Terminal-Bench Science among its reported evidence. That can justify evaluating GPT-6.1 Sol for bounded scientific computing, analysis support or tool-mediated research tasks. It does not establish that the model is a reliable scientific authority, that its calculations are correct in every domain, or that generated conclusions have been peer reviewed.

A scientific evaluation should separate procedural execution from scientific validity. Review whether commands ran, files were produced and calculations were reproducible, but also ask a qualified domain specialist to inspect assumptions, units, data exclusions, statistical methods and interpretation. A technically successful run can still answer the wrong question or apply an unsuitable method.

Recommended scientific workflow: preserve the input data, prompt, model identifier, reasoning setting, tools, software versions, generated code, intermediate artefacts and reviewer corrections. Re-run deterministic calculations independently where possible. For experimental, clinical or safety-relevant conclusions, use established review and validation procedures; a model result should not replace expert judgement, ethics review or required regulatory processes.

Factuality findings do not mean hallucinations have been eliminated

OpenAI reports factuality evidence for GPT-6.1 Sol, but the cited material does not support a claim that hallucinations have been eliminated. Factuality performance varies with topic, source availability, ambiguity, date sensitivity and whether the model is required to cite or retrieve evidence. A favourable result on an evaluation set cannot ensure correctness on proprietary records, breaking news, obscure subjects or questions with disputed premises.

Production tests should distinguish at least four failure classes: incorrect statements, unsupported statements, omitted material facts and false certainty. Checking only direct factual errors can miss answers that are technically accurate but materially incomplete. Reviewers should also record citation mismatch, where a source exists but does not support the associated claim.

Recommended factuality rule: for any output used in a decision or external publication, require claim-level verification against an authoritative and current source. If no suitable source is available, the output should be marked unverified rather than converted into a confident statement. Retrieval or file access can improve access to evidence, but neither guarantees that the model will select, interpret or cite that evidence correctly.

Decision rule: Treat OpenAI’s factuality reporting as a reason to include GPT-6.1 Sol in a controlled evaluation, not as permission to remove source checking or accountable review.

How to interpret the deployment-safety addendum

OpenAI published an addendum to the GPT-6 Astra System Card for GPT-6.1 Sol on 29 September 2026. A system card is vendor documentation describing evaluated risks, methods, safeguards and findings. It can inform a risk assessment, but it is not an independent certification, a guarantee that the model cannot enable harm, or proof that a particular deployment complies with law, regulation, contract or internal policy.

OpenAI notes that some safety evaluations use difficult or failure-eliciting prompts. This is useful for probing model behaviour under pressure, but it also limits straightforward comparison with typical use. A failure rate under an adversarial test should not automatically be interpreted as an expected production incident rate. Conversely, strong results in a controlled safety evaluation do not show that misuse, prompt injection, data exposure or unsafe tool use is impossible in a deployed system.

Risk categorisations describe OpenAI’s assessment under its stated framework and evidence. An adopting organisation still needs to assess the complete application: prompts, retrieved documents, connected systems, tool permissions, user population, monitoring, human approvals and the consequences of error. A low-risk drafting assistant and an agent able to modify production records are not equivalent deployments merely because they use the same model identifier.

Safeguards must be layered around permissions and consequences

Model-level safeguards are only one control layer. Deployment controls should include user identity and permission management, least-privilege tool permissions, data minimisation, environment separation, destination restrictions, approval gates, audit logging, output review and incident handling. Workspace role-based access control can govern who may use capabilities, but it does not by itself validate every output or make every permitted action appropriate.

Teams should assume that hostile or misleading instructions may appear inside documents, web pages, code, tickets or tool responses. Where the model can act on external content, treat that content as untrusted input. Separate instructions from data, restrict available tools, require confirmation for consequential operations and test whether the model follows embedded instructions that conflict with the authorised task.

Credentials should not be placed in prompts, test fixtures or logs. Use the organisation’s approved secret-management and short-lived credential mechanisms, and provide only the minimum scope needed for the test. Evaluation records should use synthetic or redacted examples unless production data is necessary, authorised and handled under applicable policy.

Monitoring should capture both successful and blocked behaviour. If logs retain only final answers, investigators may be unable to determine whether the model attempted an unauthorised tool call, received a misleading result or recovered from an error by taking an unintended path. Logging itself creates privacy and security obligations, so access, retention and redaction must be defined before rollout.

Red-team evidence should be mapped to the intended deployment

OpenAI’s safety evidence can help teams identify categories for internal testing, but the organisation should add risks created by its own data and tools. A software agent may face source-code prompt injection and dependency confusion. A document assistant may expose confidential content through summaries. A computer-use agent may act on the wrong account or submit a form twice. These are application-level risks that a general model evaluation cannot settle.

Recommended red-team plan: create scenarios for conflicting instructions, malicious documents, excessive data requests, identity confusion, repeated actions, permission escalation, unavailable tools and requests to bypass review. Define the expected safe behaviour before running the test. A result should be marked as a failure when the model reaches the desired business outcome through a prohibited method, not only when the final output is wrong.

Do not allow the same person to design all tests, operate the model and approve deployment for a high-impact workflow. Independent internal review can uncover optimistic scoring, omitted edge cases and controls that exist in documentation but not in the actual system. Where legal, security, privacy or safety duties apply, involve the relevant specialists before production access is granted.

A decision framework for teams considering GPT-6.1 Sol

The appropriate question is not whether GPT-6.1 Sol performs best in OpenAI’s published tables. The operational question is whether a particular configuration meets a defined quality and risk threshold for a named workflow at an acceptable total cost. The answer may differ between ChatGPT Work, Codex and an API application because each surface can provide different instructions, tools, permissions and user controls.

Candidate workflow Primary evidence to collect Required human gate Reason to pause
Repository investigation and patch drafting Accepted diagnosis, patch scope, tests, security review and tool trace Authorised engineer reviews and approves every change Unexplained edits, invented interfaces or unreliable stopping behaviour
Professional document analysis Claim-level source support, coverage, contradictions and omissions Qualified owner validates evidence and makes the decision Unsupported conclusions, silent conflict resolution or weak traceability
Computer-use automation Action logs, state verification, recovery behaviour and approval use Confirmation immediately before consequential actions Ambiguous target selection, duplicate actions or permission overreach
Scientific analysis support Reproducible calculations, assumptions, units and specialist review Domain expert approves methods and interpretation Irreproducible results, unsuitable methods or unsupported causal claims
Externally published factual content Current authoritative sources and claim-level verification Editor or accountable subject specialist signs off Source mismatch, date-sensitive uncertainty or fabricated support

Who should participate in testing

  • Workflow owner: defines the task, accepted outcome and consequences of failure.
  • Representative users: test realistic inputs, ambiguity and operational constraints rather than demonstration prompts.
  • Subject specialist: judges substantive correctness where generic reviewers cannot.
  • Security and privacy reviewers: assess permissions, data flows, logs, external content and incident handling.
  • Platform or engineering owner: records the model identifier, configuration, tools, failure handling and cost telemetry.
  • Authorised decision-maker: accepts residual risk, approves deployment scope and owns rollback criteria.

For low-risk drafting, some roles may be performed by the same person if organisational policy permits. For consequential or tool-enabled systems, role separation provides a check against optimistic interpretation. Access to GPT-6.1 Sol should also be verified on the intended surface: OpenAI’s documentation excludes regular ChatGPT conversations, and rollout in ChatGPT Work and Codex depends on plan, workspace settings and rollout status.

Set acceptance and rollback rules before comparing outputs

Recommended evaluation procedure: first define the baseline currently used by the team, including human-only processing where that is the actual alternative. Second, create a representative holdout set that includes ordinary, difficult, ambiguous and prohibited-action cases. Third, run the candidate configuration without changing the scoring rules after seeing results. Fourth, conduct blinded human review where feasible. Finally, compare quality, end-to-end latency, token use, tool behaviour, retries, review burden and failure severity.

Aggregate accuracy alone is insufficient. A useful scorecard should distinguish acceptable outputs, correctable drafts, unsafe outputs and failures requiring complete rework. Weighting must reflect the workflow: a minor formatting error should not count the same as disclosure of confidential information or an unauthorised account change. Thresholds are organisational risk decisions, not properties supplied by OpenAI’s benchmark tables.

Define rollback triggers before production use. Examples include a material increase in high-severity errors, repeated attempts to exceed permissions, unexplained cost growth, tool-call instability, loss of source traceability or a model/documentation change that invalidates the tested configuration. Rollback may mean returning to a previous model, disabling tools, narrowing the user group or restoring manual processing.

Record the date of every decision because model documentation, product rollout and workspace access can change. The evaluation record should identify gpt-6.1-sol where the API is used, the product surface for Work or Codex tests, reasoning configuration, available tools, context supplied, test-set version, reviewers and unresolved limitations. Without that record, a later re-test may not be comparable.

Adopt, restrict or defer

Adopt for a bounded workflow when representative testing meets predefined quality criteria, severe failures remain within the organisation’s accepted tolerance, permissions are no broader than necessary, human review is practical and rollback has been tested. Adoption should apply to the evaluated configuration and task class, not to every use of GPT-6.1 Sol.

Restrict to assisted use when the model produces useful drafts but substantive errors, weak source handling or inconsistent tool behaviour require regular correction. In that case, disable consequential actions, preserve human execution authority and measure whether review effort still makes the workflow worthwhile.

Defer deployment when access cannot be governed, representative data cannot be tested lawfully, high-severity failures lack reliable controls, outputs cannot be independently checked, or the organisation cannot monitor and reverse actions. OpenAI’s favourable evaluations do not override those conditions.

The defensible conclusion from the published evidence is therefore limited but useful: OpenAI has supplied reasons to evaluate GPT-6.1 Sol for coding, professional work, computer use and scientific tasks, alongside a safety addendum describing its own testing and safeguards. Each adopter must still establish suitability through representative evaluation, conservative permissions, human-led consequential decisions and documented rollout controls.

Bounded pilot release record dated 30 September 2026

Use this record only after the deployment-verification checklist has passed. It defines the operating envelope for a limited release and the evidence needed to decide whether to expand, pause or roll back. Record the named owner, participating users, permitted repositories or document classes, enabled tools, action permissions, financial limit, elapsed-time limit and escalation contact. Preserve the date and source references because rollout status, workspace permissions, model specifications and commercial conditions can change after the pilot begins. The record should also state which production systems remain out of scope, how unsuccessful tasks return to a manual path and who may safely authorise a broader release.

  1. Test permissions with low-consequence fixtures. For Codex or tool-enabled API workflows, use disposable repositories, synthetic records and non-production systems during early testing. Verify what the model can read, write, execute and transmit. Do not expose production credentials merely to discover whether an access control works.

  2. Define operational limits. Set maximum input size, output size, reasoning effort, tool-call count, retry count, elapsed time and spend per request or job where the chosen surface permits. The documented 1,050,000-token context window is a capacity limit, not a recommendation to send every available file or a guarantee that all details will be used correctly.

  3. Create escalation and fallback paths. Specify what happens when the model refuses, times out, exceeds a budget, produces an incomplete answer, requests an unavailable tool or disagrees with a deterministic control. Preserve a previously approved model or manual procedure where continuity requirements justify it.

  4. Run a bounded pilot. Limit the first deployment by users, repositories, document classes, tools and action permissions. Compare pilot evidence with the pre-agreed thresholds. Expand only after named owners accept the residual risks and support arrangements.

  5. Preserve a pilot evidence packet. Keep the approved configuration, dated test results, accepted and rejected cases, tool and permission settings, measured cost, incident notes, reviewer decisions and rollback outcome together. The packet should let a later reviewer reconstruct why the release was expanded, restricted or stopped without relying on memory or screenshots.

Decision responsibilities by role

The following table is a recommended responsibility map, not an OpenAI-prescribed governance model. Adapt it to the organisation’s accountability structure, but avoid assigning access, cost, quality and risk approval to a single model owner without independent review.

Role Primary decision Evidence to request Reason to pause
Workspace administrator Whether eligible users can access GPT-6.1 Sol in ChatGPT Work or Codex under current workspace settings Dated account check, eligible user group, workspace RBAC settings and pilot scope The team has inferred availability from a paid plan, or regular ChatGPT conversations are being treated as equivalent to Work or Codex
Engineering lead Whether gpt-6.1-sol satisfies the target API or Codex workflow Representative tests, failure analysis, tool traces, token usage, latency observations and rollback design The proposal relies only on OpenAI benchmark tables or on listed feature support without workflow testing
Product owner Whether the proposed capability improves the defined user task within accepted failure limits Task-level acceptance criteria, baseline comparison, user-review findings and escalation design The intended outcome is vague, or success is defined as model availability rather than a measurable task result
Finance or financial operations owner Whether measured usage fits the approved budget and commercial arrangement Input, cached-input and output token volumes; request counts; tool and retry costs; processing mode; regional conditions; sensitivity ranges A unit token rate has been presented as a fixed task price or universal saving
Security reviewer Whether data access, tools, credentials and action permissions are acceptably constrained Data-flow diagram, threat model, tool allowlist, secret-handling design, logs, approval gates and incident procedure The system card is being used as a substitute for deployment-specific security assessment
Privacy or data-governance owner Whether proposed inputs are authorised, minimised and handled under applicable organisational controls Data inventory, purpose, retention rules, account and workspace configuration, access list and deletion procedure Teams plan to submit sensitive data before confirming the applicable offering, settings and internal approval
Legal or compliance reviewer Whether the proposed use and contractual arrangement meet organisation-specific obligations Current terms, procurement records, processing details, human-control design and jurisdiction-specific analysis OpenAI’s safeguards or risk categorisations are being represented as legal compliance or assurance
Quality or evaluation lead Whether the evaluation supports adoption for the bounded use case Versioned test set, scoring rubric, reviewer agreement, error taxonomy, results by risk class and reproducible configuration Only aggregate averages are available, or vendor-published results are being treated as independently replicated evidence
Operations or service owner Whether monitoring, support, fallback and change control are adequate Runbook, alert criteria, usage monitoring, escalation contacts, rollback test and revalidation schedule No owner exists for incomplete responses, access changes, unexpected costs or documentation changes
Authorised end-user representative Whether outputs are useful and reviewable in the real task Blind or controlled task review, correction effort, missing-context reports and usability findings The pilot excludes the people responsible for checking or acting on the output

Cost worksheet for a measured pilot

The method below uses OpenAI’s published standard token rates without inventing a monthly bill or completed-task price. Build the worksheet from pilot telemetry, not estimates copied from a benchmark prompt. Keep each pricing category separate because cached input and output have materially different listed rates.

1. Collect usage by workload class

Create one row for each materially different task, such as repository analysis, code modification, document synthesis or image-assisted review. For every row, record successful requests, failed requests, retries, uncached input tokens, cached-input tokens, output tokens, reasoning setting, context size, tools invoked and processing mode. Separate interactive traffic from deferred processing if different commercial conditions may apply.

Worksheet field Entry method Control question
Uncached input tokens Sum measured input tokens charged at the applicable standard input category Were repeated prefixes actually eligible for cached-input treatment?
Cached-input tokens Sum only usage reported in the applicable cached category Have cache writes or other cache conditions been checked separately?
Output tokens Sum model-generated output, including output associated with retries where charged Does the measurement include verbose failures and abandoned runs?
Tool and ancillary charges Enter applicable charges from current documentation or the organisation’s agreement Has the team incorrectly assumed tools are covered by text-token rates?
Non-standard processing adjustments Apply only documented conditions for the selected mode, region or context band Is the standard rate being used for traffic subject to different terms?
Retries and failed work Retain the actual usage rather than counting only accepted outputs Does the cost-per-accepted-task calculation include unsuccessful attempts?
Human review Record review time as an internal operating input, separately from OpenAI charges Would a lower token bill still create more total work?

2. Apply the published standard token formula

For traffic to which the documented standard rates apply, calculate the text-token component as follows:

standard_token_cost =
  (uncached_input_tokens / 1,000,000 × $2)
+ (cached_input_tokens / 1,000,000 × $0.10)
+ (output_tokens / 1,000,000 × $10)

Add any applicable non-token charges or documented adjustments only after checking the current model reference and the organisation’s commercial terms. Do not assume that cache writes, long-context traffic, Fast mode, Batch processing or Flex processing, regional processing or tool use share the three standard rates above.

3. Calculate a range rather than one optimistic forecast

Recommended method: prepare low, expected and high usage cases using observed variation in tokens, retries and acceptance rates. The high case should include difficult inputs, tool failures and outputs that require regeneration. This exposes whether the budget remains acceptable when production traffic differs from the median pilot request.

For operational comparison, divide the fully measured pilot cost by the number of outputs that passed the predefined quality gate, not by total requests. Label the result “observed cost per accepted pilot task” and include the sample period, task mix and acceptance rule. It remains a local planning measure, not a universal price or a prediction for another workload.

4. Reconcile invoices and telemetry

Before wider rollout, compare the worksheet with available billing records and investigate unexplained differences. Check model identifiers, processing modes, regional settings, tool use, retries and traffic outside the pilot. Do not force the invoice total to match a simplified formula by silently reclassifying usage.

Material risks and required controls

  • Surface confusion: a team may assume paid ChatGPT access includes GPT-6.1 Sol in ordinary conversations. Control this by naming the surface in requirements, training and support records, then verifying access in the intended workspace.
  • Rollout uncertainty: announcement language and staged release notes can produce different expectations. Use a dated account check and define a fallback if access has not reached the relevant users.
  • Cost extrapolation: published token rates may be converted into an unsupported claim about task savings. Require measured token mix, retries, tool use, processing conditions and human review before approving a forecast.
  • Benchmark overreach: OpenAI reports favourable results against GPT-6 Sol and selected comparators, but evaluated tasks and conditions may not match production. Label these results as vendor-published and run a representative local evaluation.
  • Large-context misuse: the 1,050,000-token context window may encourage indiscriminate data submission. Minimise inputs, test retrieval quality and verify that critical evidence is actually reflected in the output.
  • Tool-enabled consequences: an incorrect tool call can change systems rather than merely produce flawed text. Use least privilege, bounded tools, dry runs, approval gates and deterministic validation for consequential operations.
  • Security and misuse residuals: the deployment-safety addendum documents risk categorisations and safeguards, but does not prove that harmful use is impossible. Apply organisation-specific threat modelling, monitoring and incident response.
  • Compliance assumptions: neither a model specification nor a system card establishes compliance for a particular deployment. Keep procurement, privacy, legal and security approval human-led and tied to the actual data, account, region, contract and workflow.
  • Configuration drift: model documentation, workspace access and product behaviour may change after approval. Monitor primary documentation and re-run critical tests after material changes.
  • Unclear accountability: teams may treat the model owner as responsible for every downstream decision. Name separate owners for access, data, security, quality, budget, operations and final human actions.

Questions teams are likely to ask

Can every paid ChatGPT user use GPT-6.1 Sol?

No universal access claim is supported by the cited documentation. OpenAI announced paid-plan availability in ChatGPT Work and Codex from 29 September 2026, while its release notes describe a staged rollout starting with Pro. Access also depends on workspace settings and permissions. Check the actual account and workspace.

Is GPT-6.1 Sol available in ordinary ChatGPT conversations?

OpenAI’s cited documentation says it is not available in regular ChatGPT conversations. ChatGPT Work, Codex and regular ChatGPT conversations are distinct product surfaces for this release and should not be described interchangeably.

What model name should an API developer configure?

The documented model identifier is gpt-6.1-sol. Record that exact identifier in test evidence and deployment configuration. Recheck the model reference before release rather than relying on a copied identifier from an earlier integration.

Does the $2 input rate mean a task costs $2?

No. The figure is a standard rate per million input tokens, not a task price. OpenAI also lists $0.10 per million cached-input tokens and $10 per million output tokens. A task’s charge depends on its token mix and may also be affected by caching conditions, long context, reasoning, tools, processing mode, region and commercial terms.

Does cached input make every repeated request cheaper?

No such universal conclusion follows from the rate card. Confirm that the usage qualifies for cached-input treatment and account separately for any applicable cache-write conditions. Measure the reported categories rather than estimating cache benefits from visually similar prompts.

Can the full context window be treated as reliable working memory?

No. OpenAI documents a 1,050,000-token context window, which describes capacity rather than guaranteed retrieval, reasoning or factual accuracy across every token. Test long documents for omission, source confusion and misplaced emphasis, and avoid including material that is unnecessary or unauthorised.

Does GPT-6.1 Sol process audio or video directly?

The cited model reference lists text and image input and text output, with no audio or video modalities. If audio or video is part of the workflow, evaluate and govern the separate conversion stage rather than attributing that capability to GPT-6.1 Sol.

Do OpenAI’s benchmark results prove that GPT-6.1 Sol will outperform the current model?

No. The announcement and deployment-safety materials report OpenAI’s evaluations under defined conditions. They are not independent replication or a guarantee for a particular repository, desktop environment, document set or professional task. Use them to form hypotheses and test cases, not to skip a local comparison.

Does the system card make the deployment safe or compliant?

No. The system card and addendum provide relevant information about evaluated risks and safeguards, but they are not assurance that the model cannot enable harm and are not a compliance certification for the reader’s deployment. Security, privacy, legal and operational reviews must address the actual implementation.

Should teams replace GPT-6 Sol immediately?

The sources support describing GPT-6.1 Sol as OpenAI’s upgrade to GPT-6 Sol, not an instruction to replace every production configuration immediately. Run a bounded evaluation, compare accepted-task quality and total operating cost, verify access and preserve rollback until the new configuration passes the organisation’s gates.

What should be recorded for audit or change control?

At minimum, record the decision date, product surface, account or workspace, model identifier, documentation dates, configuration, reasoning setting, tool permissions, evaluation set, acceptance rules, error findings, token usage, cost assumptions, approvers, residual risks, rollback procedure and next review date. Retain enough evidence to reproduce the decision without storing unnecessary sensitive prompts or credentials.

Bottom line

As of 30 September 2026, GPT-6.1 Sol is a separately documented model for ChatGPT Work, Codex and the API, with staged access that must be verified in the relevant account or workspace. It should not be reported as available in regular ChatGPT conversations. API teams should use the documented gpt-6.1-sol identifier and validate its listed text-and-image input, text output, context and tool-calling capabilities against the intended workflow.

The published rates provide a basis for a measured worksheet, not a universal saving or fixed task price. Likewise, OpenAI’s benchmark, factuality, alignment and safety findings are vendor-published evidence bounded by their evaluated conditions. A defensible adoption decision combines live access verification, representative testing, measured token and tool usage, least-privilege controls, human approval for consequential actions and scheduled revalidation of the primary documentation.

Evidence boundary: OpenAI reports the GPT-6.1 Sol results cited here. They are OpenAI-reported, vendor-published results, not an independent evaluation. As of 30 September 2026, OpenAI documentation lists GPT-6.1 Sol for ChatGPT Work, Codex and the API, and excludes it from regular ChatGPT conversations. The documented API model identifier is gpt-6.1-sol; the listed standard rates are $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens, with documented qualifiers for cache writes, long context, Fast, Batch, Flex and regional processing. The documented context window is 1,050,000 tokens.

A surface-and-identifier control matrix

Before a team evaluates quality or cost, it should resolve two separate questions: where the model is supposed to appear and which identifier the workflow actually uses. A visible model option in ChatGPT Work or Codex does not prove that an API integration is configured for the same model. Likewise, successful API access to gpt-6.1-sol does not imply availability in regular ChatGPT conversations.

Surface What to verify Evidence to retain Do not infer
ChatGPT Work The option is visible to the intended user, under the intended paid plan and workspace. Date, account type, workspace, visible model label, and applicable workspace setting. That every paid user has access or that regular ChatGPT conversations include the model.
Codex The intended Codex workflow can select or use the model under current rollout and workspace controls. Date, user role, workspace, selected model, and relevant permission state. That availability in Codex transfers to ChatGPT conversations or to an unrelated workspace.
API The request uses the documented gpt-6.1-sol identifier and succeeds for the relevant project and commercial arrangement. Request configuration, returned model metadata where available, endpoint, date, project, and error details. That a user interface (UI)The controls and visual surfaces through which a person interacts with software. Open glossary entry entitlement grants API access, or that standard rates describe every processing mode.
Regular ChatGPT conversations The cited documentation explicitly excludes GPT-6.1 Sol from regular conversations. The dated documentation used for the access decision. That a similar name, paid subscription, or staged rollout changes this documented boundary.

Treat each row as an independent control. If one surface works and another does not, that can be consistent with the documented product boundary rather than evidence of a model-wide outage. Because rollout and documentation can change, record the date of every access check instead of turning a point-in-time result into a permanent availability statement.

Failure analysis: diagnose the layer before changing the prompt

When GPT-6.1 Sol does not behave as expected, immediately rewriting the prompt can conceal the real cause. Use a layered procedure so that access, configuration, cost, tool, and output failures are not treated as one problem.

  1. Freeze the evidence.

    Retain the timestamp, product surface, workspace or API project, model identifier, endpoint, reasoning setting, processing mode, input size, tool configuration, response status, and relevant logs. Remove or protect sensitive content according to the organisation’s controls, but preserve enough information to reproduce the condition.

  2. Classify the failure layer.

    • Access: the model is absent, denied, or unavailable to the account or workspace.
    • Configuration: the wrong identifier, endpoint, modality, reasoning setting, or tool declaration was used.
    • Input: the source material was incomplete, contradictory, poorly selected, or larger than the application handled reliably.
    • Model output: the response was unsupported, incomplete, inaccurate, or inconsistent with the acceptance rule.
    • Tool execution: a tool failed, returned unexpected data, lacked authorisation, or attempted an operation outside policy.
    • Economics: retries, output volume, cache misses, long context, tools, or processing choices made completed work costlier than forecast.
  3. Test the smallest discriminating change.

    Change one variable where practical. For an access problem, verify the surface and workspace before altering prompts. For a cost problem, inspect token mix and retries before claiming that the published rate is wrong. For a quality problem, rerun the same case under the recorded configuration before changing both instructions and source material.

  4. Compare against the acceptance rule.

    Describe the failure in operational terms: which requirement was missed, what evidence demonstrates the miss, and what consequence followed. “The model was worse” is not sufficiently specific for corrective action.

  5. Choose containment.

    Possible dispositions are to retry under a documented condition, route the case to a person, disable a tool, narrow the workload, revert the model configuration, or pause adoption. The appropriate response depends on the consequence of the failure.

  6. Update the test set.

    Add a sanitised version of the incident to regression testing when authorised. A corrected single case is not evidence that the broader failure class has been resolved.

Worked hypothetical examples

Example 1: A paid user cannot find GPT-6.1 Sol

Hypothetical example: A paid user opens an ordinary ChatGPT conversation and expects to select GPT-6.1 Sol because the announcement refers to paid-plan availability. The option is not present.

The first decision is not to troubleshoot the user’s prompt. The cited documentation separates regular ChatGPT conversations from ChatGPT Work and Codex and explicitly says GPT-6.1 Sol is unavailable in regular conversations. The team next checks whether the user is entering the intended Work or Codex surface, whether the relevant workspace permits access, and whether the staged rollout has reached that account. Release notes beginning rollout with Pro do not establish immediate availability for every otherwise eligible user.

The incident record should therefore say “surface or rollout access not verified,” not “GPT-6.1 Sol is down.” If access is required for a scheduled pilot, the team should defer that user’s participation or use an already approved alternative rather than representing regular ChatGPT conversations as an equivalent route.

Example 2: A low token-rate estimate understates completed-task cost

Hypothetical example: A team estimates a document workflow using only the listed $2-per-million input rate. Its pilot then produces longer outputs, repeated retries, tool activity, and fewer cache hits than assumed.

The published rate has not become a fixed price promise. The estimate omitted the separately listed $10-per-million output rate and did not model cached versus uncached input correctly. It also excluded retry and tool behaviour. The team should reconstruct each accepted task from measured input, cached input, output, and other applicable charges or terms. It should then report a range across observed task classes rather than claiming that the model costs a fixed amount per document.

Example 3: A large context window encourages an unsafe review shortcut

Hypothetical example: An analyst proposes placing an entire source pack into one request because the documentation lists a 1,050,000-token context window. The proposed workflow would accept the resulting summary without checking source support.

The capacity listing establishes a documented limit, not reliable comprehension of every detail or production suitability. A safer evaluation defines required source citations or evidence locations, tests omission and contradiction cases, and has reviewers compare consequential claims with the source pack. If the workflow cannot identify whether required evidence was considered, the large context capacity does not cure that control gap.

Example 4: A favourable benchmark is used as an adoption argument

Hypothetical example: A project owner argues that OpenAI’s reported coding and computer-use results remove the need for a local pilot.

The decision record should distinguish vendor-published evaluations from independent proof. The evaluated tasks and conditions may differ from the organisation’s repositories, tools, permissions, and failure costs. The team can use the reported results to justify testing GPT-6.1 Sol as a candidate, but adoption should still depend on representative local cases, action-level controls, and predefined acceptance and rollback rules.

Additional frequently asked questions

Is a model label in the interface enough evidence for change control?

No. Record the surface, workspace, account context, date, and applicable permissions. For API use, also record the configured model identifier and endpoint. A label observed in one product surface does not establish availability or configuration in another.

Should an API integration silently fall back if gpt-6.1-sol is unavailable?

Only if the organisation has deliberately designed, tested, and approved that behaviour. A silent fallback can change quality, cost, tool behaviour, or evidence records. At minimum, the application should make the selected model observable and apply the acceptance rules defined for that configuration.

Does support for image input mean every image-based workflow is suitable?

No. The specification documents image input, but feature support does not guarantee successful interpretation of every image or suitability for a consequential workflow. Representative source material and human verification remain necessary where errors matter.

Can a team approve all professional-work use after one successful pilot?

No general approval follows from one workload. Approval should identify the tested task, source types, tools, reasoning configuration, review rules, and consequence limits. Materially different work should receive its own evaluation.

What should trigger an immediate pause?

Teams should define their own triggers before launch. Candidates include loss of verified access, an unapproved model or mode change, unauthorised tool behaviour, failure of a critical acceptance rule, unexplained cost variance, or inability to reconstruct consequential outputs. The exact threshold is a risk decision, not a universal property of GPT-6.1 Sol.

How often should the access and documentation record be refreshed?

Refresh it whenever the rollout state, workspace configuration, model documentation, processing mode, commercial arrangement, or application configuration changes. Because the documentation and rollout can change, a dated record should not be treated as permanently current.

Does a system-card risk category authorise deployment?

No. The addendum provides OpenAI’s evaluation and safeguard context. It does not grant organisational approval, establish legal or regulatory compliance, or show that the model cannot enable harm. Deployment authority and controls remain separate decisions.

What is the minimum defensible outcome of an unsuccessful pilot?

A useful unsuccessful pilot identifies the tested configuration, representative cases, failure classes, cost observations, containment decision, and conditions for reconsideration. It should not be converted into a claim that the model fails every workflow, just as a successful pilot should not be generalised beyond its tested scope.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library →

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this