OpenAI Launches GPT-6 Sol and Luna: Half-Price APIs, ChatGPT Work and Codex Rollout, and New Caching Controls


OpenAI adds two lower-cost GPT-6 models, but keeps Astra at the top
OpenAI announced GPT-6 Sol and GPT-6 Luna on September 22, 2026, expanding the GPT-6 family with two faster, less expensive models positioned below GPT-6 Astra. In OpenAI’s launch framing, Astra remains the company’s strongest overall model for the most demanding work, while Sol and Luna offer different capability-and-cost tradeoffs for professional tasks, coding workflows, and agent-style use cases. The practical takeaway is not that teams should replace Astra everywhere; it is that OpenAI now exposes more price-performance tiers inside the GPT-6 line, and each tier needs separate evaluation before production routing changes.
The exact API model identifiers are gpt-6-sol and gpt-6-luna. Those names matter operationally because routing rules, observability dashboards, evaluation logs, billing attribution, and rollback procedures should use the model ID actually sent to the API rather than a marketing label such as “Sol,” “Luna,” or “GPT-6.” OpenAI’s model-reference pages identify Sol and Luna as GPT-6 family models with text and image inputs, text output, prompt caching, Responses API support, Chat Completions support, structured outputs, streaming, Batch, and a set of supported tools; the same pages also list unsupported endpoint families such as Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, and legacy Completions for these models.
OpenAI’s published Standard API list prices put Sol at $2 per million input tokens and $10 per million output tokens, and Luna at $0.10 per million input tokens and $0.50 per million output tokens. OpenAI describes those as 50% cheaper than GPT-5.6 promotional prices, but that statement should not be converted into a promise that a real application’s bill will fall by 50%. Real cost depends on the full request shape: uncached input tokens, cache writes, cached reads, output tokens, invisible reasoning tokens, tool calls, retries, processing mode, regional processing premiums where available, and the long-context multipliers that apply to requests above 272,000 input tokens.
Availability is split across product surfaces, which is one of the most important launch details for administrators and developers. OpenAI says Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. OpenAI also says Free and Go users can access Luna in the desktop application. At the cited launch/update time, OpenAI’s launch page and Help Center state that Sol and Luna are not yet available in ordinary Chat conversations. That means a user seeing Luna in a desktop Work or Codex context should not assume the same model is available in normal Chat, web/mobile Chat, API usage, or another workspace.
This article covers the GPT-6 Astra launch across ChatGPT, Codex, the API, Azure, and AWS Bedrock, including staged availability, benchmarks, and immediate product implications. The GPT-6 Astra Launches in ChatGPT, Codex, the API, Azure, and AWS Bedrock: Availability, Benchmarks, and What Changes Now article is a focused companion for GPT-6 Astra Launch Analysis because it is the most exact background link for a marker referring to GPT-6 Astra launch analysis and helps contextualize Sol and Luna against the earlier Astra rollout.
The launch also arrives alongside OpenAI’s “better prompt caching for GPT-6” announcement, which describes higher cache hit rates by default and new developer controls such as a Prompt Caching Dashboard, diagnostics, explicit breakpoints, prewarming, and append-only prompt patterns. For teams already spending heavily on repeated system prompts, tool schemas, long instruction blocks, or stable document context, the caching update may be as financially important as the new model list prices. However, prompt caching is prefix reuse, not semantic memory, quality validation, or authorization; it reduces repeated computation only when the rendered prefix and cache-sensitive settings remain compatible.
Launch facts versus nonclaims
The launch creates several easy-to-misread headlines: “half price,” “available in ChatGPT,” “better caching,” and “stronger coding.” Each is directionally meaningful, but each also has a boundary in OpenAI’s source material. Developers, founders, procurement teams, and security reviewers should separate what OpenAI actually announced from conclusions that require local proof.
| Topic | Launch fact attributed to OpenAI | What that fact does not establish | Operational decision rule |
|---|---|---|---|
| Family positioning | Sol and Luna are faster, more affordable GPT-6 family members; Astra remains OpenAI’s strongest model overall. | It does not prove that Sol or Luna can replace Astra for a specific enterprise workflow, regulated process, coding pipeline, or agent loop. | Keep Astra available for high-risk or hardest tasks until local evaluations show whether Sol or Luna can meet the same acceptance criteria with the same controls. |
| API model IDs | The API identifiers are gpt-6-sol and gpt-6-luna. |
It does not create aliases, automatic migrations, or guaranteed availability in every account, region, endpoint, or processing mode. | Record exact model IDs in logs, eval reports, cost dashboards, and rollback runbooks; validate endpoint and tool support before changing production traffic. |
| API list prices | Sol is listed at $2 input and $10 output per million tokens; Luna is listed at $0.10 input and $0.50 output per million tokens. | It does not guarantee lower total task cost, because output length, reasoning tokens, cache writes, tool fees, retries, Fast mode, regional premiums, and long-context multipliers can change the bill. | Compare models using measured cost per accepted task, not only published input-token price. |
| “50% cheaper” claim | OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional prices. | It does not mean every GPT-5.6 workload receives a 50% realized cost reduction after migration. | Run cache-normalized before-and-after tests with representative traffic and identical acceptance criteria. |
| ChatGPT Work and Codex | Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, subject to access conditions. | It does not mean every seat, role, workspace, or app surface receives identical model access immediately. | Administrators should verify plan, workspace policy, role permissions, rollout access, default model settings, and product surface separately. |
| Free and Go desktop access | OpenAI says Free and Go users can access Luna in the desktop application. | It does not imply Sol access for Free or Go users, API access, ordinary Chat access, or identical web/mobile behavior. | Document which surface is approved for which user group; avoid support instructions that treat desktop, Chat, Work, Codex, and API as interchangeable. |
| Ordinary Chat | At launch, OpenAI states that Sol and Luna are not yet available in Chat. | It does not rule out future rollout changes, and it does not clarify every account’s future timing. | Write help-desk and admin guidance using dated language: “at the cited launch/update time,” not permanent claims. |
| Benchmarks | OpenAI reports evaluation-specific improvements in professional work, factuality, coding, computer use, and collaboration style. | It does not guarantee performance in a private codebase, legal review flow, enterprise data environment, classroom, or healthcare-adjacent workflow. | Use public benchmarks as hypotheses; gate migrations on local task-specific evaluations and human review. |
| Alignment evaluations | OpenAI reports improvements over GPT-5.6 counterparts in several challenge-set alignment tests, including lower rates of misleading claims about coding work. | Challenge tests are not typical-use incident-rate estimates and do not prove that a model will never misreport progress or take unsafe action. | Keep audit logs, tool restrictions, sandboxing, review gates, and rollback procedures in place regardless of model selection. |
| Prompt caching | OpenAI says GPT-6 prompt caching can provide discounts of up to 90% on eligible cached input-token reads and supports more developer controls. | It does not guarantee a cache hit, semantic correctness, data authorization, or full-request reuse. | Measure cached_tokens, cache-write cost, read frequency, and miss reasons; do not treat caching as a privacy or governance substitute. |
Where Sol and Luna fit on OpenAI’s cost-intelligence curve
OpenAI’s model family now reads like a three-level GPT-6 ladder: Luna at the lowest published token price, Sol in the middle, and Astra at the top for the hardest work. This is best understood as a cost-intelligence curve rather than a simple ranking. A low-price model can be the correct choice for high-volume classification, summarization, extraction, formatting, or simple coding assistance; a stronger model can still be cheaper for a difficult task if it needs fewer retries, shorter prompts, less human correction, or fewer failed tool calls.
Luna’s published pricing makes it the most aggressive cost play in the launch: $0.10 per million input tokens and $0.50 per million output tokens under Standard API pricing. That makes Luna a candidate for focused, high-volume workflows where the task can be clearly specified, error consequences are low or reviewed, and acceptance tests can catch failures. Examples include generating first-pass issue summaries from approved internal records, transforming already-verified content into a structured format, drafting low-risk internal notes for human review, or triaging repetitive support categories without making customer-facing decisions automatically.
Sol sits between Luna and Astra, at $2 per million input tokens and $10 per million output tokens under Standard API pricing. OpenAI positions Sol for stronger professional work and coding use cases than a low-cost model alone would typically cover. A team might test Sol for complex codebase analysis, multi-file refactoring suggestions, tool-using internal agents, long technical document synthesis, or higher-stakes knowledge-work drafts that still receive human approval before publication, deployment, or external communication.
Astra remains the strongest overall GPT-6 model according to OpenAI, and that positioning matters for governance. The existence of cheaper Sol and Luna routes does not reduce the need for Astra in difficult, ambiguous, high-impact, or safety-sensitive workflows. If a task involves legal commitments, privileged material, regulated decisions, destructive operations, production deployments, financial transactions, medical or youth-safety implications, or external representations by the organization, model choice should be only one control among many. Authorization, least privilege, sandboxing, audit logging, retrieval verification, policy enforcement, and human approval remain mandatory.
This article explains OpenAI’s October 14 GPT-5.5 retirement notice for ChatGPT, ChatGPT Work, and Codex, with emphasis on product-surface deadlines and API boundaries. The GPT-5.5 Retirement Set for October 14: Product-Surface Deadline, Codex Replacements, and the API Boundary article is a focused companion for GPT-5.5 Retirement Boundary because it directly matches the marker’s retirement-boundary concept and is relevant to an article discussing model lineup changes and rollout impacts.
API pricing: list prices, cache prices, and the costs hidden outside the headline
The most quotable numbers in the launch are the Standard API token list prices, but they are not a complete cost model. OpenAI’s model pages list uncached input, cached input, cache write, and output rates separately. For Sol, the cited Standard API prices are $2 input, $0.20 cached input, $2.50 cache write, and $10 output per million tokens. For Luna, the cited Standard API prices are $0.10 input, $0.01 cached input, $0.125 cache write, and $0.50 output per million tokens. These rates reflect OpenAI’s documented prompt-caching economics for GPT-5.6-and-later models: cache writes cost 1.25 times the uncached input rate, while cached reads cost 0.1 times the uncached input rate.
| Model | Exact API model ID | Uncached input price | Cached input read price | Cache write price | Output price |
|---|---|---|---|---|---|
| GPT-6 Luna | gpt-6-luna |
$0.10 / 1M tokens | $0.01 / 1M tokens | $0.125 / 1M tokens | $0.50 / 1M tokens |
| GPT-6 Sol | gpt-6-sol |
$2 / 1M tokens | $0.20 / 1M tokens | $2.50 / 1M tokens | $10 / 1M tokens |
A single cache write is not automatically cheaper than ordinary input processing. If an application writes a large prefix once and never reuses it, the write costs more than a standard uncached input pass for that prefix. Caching becomes economically useful when a stable rendered prefix is reused enough times to offset the 1.25x write cost and when the rest of the request does not dominate the total bill. A support agent that reuses the same long policy manual across many requests may benefit; a one-off research prompt that rewrites its instructions and documents every time may not.
OpenAI also documents a long-context pricing threshold for these model pages: for requests above 272,000 input tokens, input and cache rates are multiplied by 2x and output rates by 1.5x for the full request. That detail is essential for enterprise retrieval, litigation-support review, large repository analysis, and classroom-scale corpus workflows, because a single oversized request can change the effective economics of a model comparison. A narrower retrieval step that sends fewer approved passages may outperform a brute-force “paste everything” strategy in both cost and quality.
Reasoning tokens add another cost dimension. OpenAI’s reasoning documentation states that reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible. Sol and Luna default to medium reasoning effort, and their model pages list supported efforts including none, low, medium, high, xhigh, and max. More reasoning effort can be valuable for difficult tasks, but it can also increase token usage, latency, and the risk that max_output_tokens produces an incomplete response before useful visible text appears. Applications should explicitly handle incomplete status rather than treating any model response as usable.
ChatGPT Work, Codex, desktop, Chat, and API are separate rollout surfaces
The launch is unusually easy to misunderstand because “available in ChatGPT” is not precise enough. OpenAI’s Help Center separates ordinary Chat from ChatGPT Work and Codex. It says Sol and Luna are models for ChatGPT Work and Codex and are not available in ordinary Chat conversations at the cited update time. It also says availability depends on plan, workspace settings, role permissions, and rollout access. For enterprise administrators, that means a model-picker setting is not the same thing as an entitlement, and a workspace default does not grant a model to a user whose role or plan lacks access.
Codex adds another distinction: OpenAI’s Help Center says Codex preserves a manually selected model, while the desktop Work/Codex picker and defaults are separate from ordinary Chat defaults. It also states that Codex CLI does not use the desktop slider. A developer who changes a desktop Work/Codex setting should not assume the same selection applies to CLI-based Codex workflows, API-backed coding agents, or ordinary Chat sessions. Support documentation should name the exact surface: Work in the desktop app, Codex, Codex CLI, ordinary Chat, or API.
Free and Go users receive a narrower launch benefit: OpenAI says they can access Luna in the desktop application. That does not imply access to Sol, access in ordinary Chat, access in browser or mobile surfaces, or API access. It also does not make ChatGPT plan usage equivalent to API billing. OpenAI’s Help Center notes that ChatGPT plan usage and API-key billing are separate systems, so procurement teams should not forecast API spend from plan allowances or treat API list prices as a substitute for seat-level ChatGPT policy.
| Surface | What OpenAI says at launch/update time | Important boundary | Admin or developer action |
|---|---|---|---|
| Ordinary Chat | Sol and Luna are not yet available in Chat according to the launch page and Help Center language cited for this report. | Do not treat Work/Codex availability as ordinary Chat availability. | Use dated release notes and avoid telling users that a model is available “in ChatGPT” without naming the surface. |
| ChatGPT Work | Sol and Luna are available for Plus, Pro, Business, Enterprise, and Edu users, subject to plan, workspace settings, role permissions, and rollout access. | A workspace starting default does not grant unavailable models to a role. | Check model access by role and workspace; document defaults separately from permissions. |
| Codex | Sol and Luna are available in Codex for specified paid and organizational plans, subject to access controls and rollout. | Codex model selection behavior is separate from ordinary Chat defaults, and Codex CLI does not use the desktop slider. | Validate Codex, desktop, and CLI behavior independently before issuing developer migration instructions. |
| Desktop app for Free and Go | Luna is available to Free and Go users in the desktop application. | This is not a general Sol entitlement and not necessarily an API or ordinary Chat entitlement. | Restrict help-center language to Luna, Free/Go, and desktop unless OpenAI updates the source policy. |
| API | The API model IDs are gpt-6-sol and gpt-6-luna, with published model-reference capabilities and prices. |
API pricing, tool support, processing modes, endpoint support, and billing are separate from ChatGPT plan usage. | Run API-specific endpoint tests, cost measurement, safety review, and staged rollout before production migration. |
What OpenAI claims about capability, and how teams should read it
OpenAI reports improvements for Sol and Luna across professional work, factuality, coding, computer use, and collaboration style. The company’s launch article includes benchmark examples such as Sol xhigh at 33.2% on AutomationBench for $0.27 per task, Sol max at 68.8% on DeepSWE 1.1, and Luna max at 66.6% on DeepSWE 1.1. Those numbers are meaningful as vendor-reported evaluation results under specified setups, but they should not be treated as a universal production ranking. OpenAI’s own launch language notes that evaluations may differ from production ChatGPT because of system prompts, tools, and other deployment differences.
For coding teams, the DeepSWE and related coding results are useful starting signals, not merge criteria. A coding agent’s local quality depends on repository conventions, build reproducibility, dependency availability, test coverage, tool permissions, branch protections, review policies, and whether the model can inspect enough relevant context without leaking secrets or overloading long-context pricing. A model that performs well on a public benchmark can still fail on a private monorepo with generated code, legacy patterns, sparse tests, or brittle deployment scripts.
For knowledge workers and educators, the factuality and collaboration-style claims should be treated with the same caution. A model’s knowledge cutoff is not a freshness guarantee, and the model pages list different cutoffs for GPT-6 family members. Sol’s published knowledge cutoff is April 20, 2026, while Luna’s is May 18, 2026. A later cutoff is not a quality ranking and does not remove the need for retrieval, dated sources, citations, and human verification when the output affects a classroom, customer, compliance record, publication, or business decision.
For legal-technology professionals, the launch does not authorize unattended legal analysis, filing, privilege calls, contract commitments, discovery decisions, or client communications. Sol or Luna may be useful for drafting, issue spotting, summarizing approved documents, or generating review checklists, but any legal conclusion, representation, production decision, or external submission requires qualified human review. Model price reductions also do not justify expanding access to privileged or confidential materials beyond approved roles, matters, regions, or retention policies.
Prompt caching becomes a bigger part of GPT-6 cost control
OpenAI’s companion prompt-caching announcement says GPT-6 prompt caching can provide discounts of up to 90% on eligible cached input-token reads and describes higher cache hit rates by default. The API documentation explains the underlying mechanism more narrowly: prompt caching reuses an unchanged rendered prefix and stores key-value tensors rather than prompt tokens. That distinction matters because caching is not a searchable memory of prior conversations, not a truth layer, and not an authorization system. A cache hit can reduce computation on a matching prefix, but it does not make the content correct, current, safe, or permitted for the next user.
For GPT-5.6 and later, OpenAI documents a minimum cacheable visible prefix of 1,024 tokens. Shared prefixes remain eligible for reuse within a 30-minute window, and the current prompt_cache_options.ttl value is 30m. The documentation describes that as a minimum eligibility period after the latest write or reuse, while noting that OpenAI may retain entries longer. Security teams should therefore avoid treating 30 minutes as a guaranteed physical-retention ceiling. Sensitive-data handling, Zero Data Retention eligibility, regional processing boundaries, logging controls, and tenant isolation still require separate governance review.
The most practical caching change for GPT-6 developers is the ability, on supported GPT-6 models, to append a configuration_update to alter reasoning effort during a conversation while preserving the earlier cacheable prefix. OpenAI warns that changing top-level reasoning settings can affect cache reuse. The same principle applies to tools, schemas, verbosity, context management, and earlier input changes: if an application rewrites the stable prefix, reorders tool definitions, changes structured output schemas, compacts earlier turns, or changes cache-sensitive settings, it may lose reuse. The recommended pattern is to keep stable instructions and tool definitions stable, then append new instructions later when semantically correct.
Tool definitions deserve special attention because agentic and coding workflows often mutate tool lists dynamically. OpenAI’s caching guidance says tool definitions, schemas, and ordering should remain stable where possible, and developers should use allowed_tools or tool_choice: none instead of removing definitions when callability changes. That is a cost and correctness practice, not a permission bypass. An application must still enforce least privilege server-side; a cheaper cached tool schema does not authorize the model to call tools that the user, role, workspace, region, or workflow should not be allowed to use.
Safety, alignment, and human approval remain part of the launch story
OpenAI’s Deployment Safety materials for GPT-6 Astra were updated on September 22, 2026 with a Sol/Luna appendix covering safety, robustness, health, hallucinations, alignment, monitorability, preparedness, and safeguards. The system-card framing is deliberately cautious: evaluation results are defined-condition results, not proof of universal behavior, zero risk, or production incident rates. The card also emphasizes remaining failures, evaluation awareness, monitorability limits, and the fact that absence of observed failures does not establish reliability across settings.
That caveat is especially important for coding and agentic work. OpenAI reports that Sol and Luna improve over GPT-5.6 counterparts in several alignment tests, including lower rates of misleading claims about coding work, but these are challenge evaluations. They should not be converted into claims that a coding agent will never overstate test success, invent implementation details, hide uncertainty, or imply a task is complete when it is not. Teams should require verifiable artifacts: test logs, diffs, build results, static analysis output, tool-call traces, review comments, and human approval before merge or deployment.
For enterprise administrators, the correct governance response is not to block every new model by default or to enable every model because it is cheaper. The defensible middle path is staged access: start with non-production evaluation, restrict tools and data, record exact model IDs and reasoning settings, compare outputs against human-labeled tasks, monitor cached-token behavior and cost, and expand access only where the model meets objective acceptance thresholds. Any workflow involving external messages, publication, account changes, payments, purchases, bookings, permission changes, destructive actions, legal commitments, regulated decisions, or production code deployment should retain explicit human approval.
Operational warning: A lower-cost model route never authorizes broader action scope. If Luna or Sol replaces a more expensive model in an agent, the agent’s permissions, data boundaries, review gates, audit logs, sandbox restrictions, and rollback path must remain at least as strong as before.
Sol and Luna pricing only looks simple until you price the whole request
OpenAI’s published Standard API prices put GPT-6 Sol and GPT-6 Luna into very different cost bands, even though both are lower-cost members of the GPT-6 family below Astra. For GPT-6 Sol, the model page lists $2 per one million input tokens, $0.20 per one million cached input tokens, $2.50 per one million cache-write tokens, and $10 per one million output tokens. For GPT-6 Luna, the listed prices are $0.10 per one million input tokens, $0.01 per one million cached input tokens, $0.125 per one million cache-write tokens, and $0.50 per one million output tokens. Those are API token prices, not end-to-end task prices, and the difference matters for engineering, procurement, and cost-allocation decisions.
OpenAI’s launch article describes both Sol and Luna as 50% cheaper than their GPT-5.6 promotional prices. That comparison is useful as a launch headline, but it should not be converted into a blanket “our bill will fall by 50%” forecast. A workload that asks for longer answers, turns on higher reasoning effort, calls tools more often, retries failures, crosses the long-context threshold, or writes large prompt caches can spend more than a workload that uses an older model with shorter context and fewer output tokens. The correct comparison is per workflow: same input corpus, same tool authority, same success criteria, same review burden, same retry policy, and a cost calculation that separates uncached input, cache writes, cached reads, output, reasoning tokens, tool charges, regional premiums, and processing mode.
| Model | Uncached input | Cached input read | Cache write | Output | Practical interpretation |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.10 / 1M tokens | $0.01 / 1M tokens | $0.125 / 1M tokens | $0.50 / 1M tokens | Lowest listed GPT-6-family token price among the cited Sol/Luna/Astra model pages; best evaluated for high-volume, focused tasks before assuming it can replace stronger routes. |
| GPT-6 Sol | $2 / 1M tokens | $0.20 / 1M tokens | $2.50 / 1M tokens | $10 / 1M tokens | Higher price than Luna, lower than Astra, and positioned by OpenAI for more complex coding and agentic work; still requires application-specific evaluation. |
| GPT-6 Astra | $10 / 1M tokens | $1 / 1M tokens | $12.50 / 1M tokens | $50 / 1M tokens | OpenAI continues to position Astra as the strongest overall GPT-6 model; included here as a pricing anchor, not as a reason to route every task to the most expensive model. |
The cache-write price is often the first number teams miss. OpenAI’s prompt-caching documentation for GPT-5.6 and later states that cache writes cost 1.25 times the uncached input rate and cached reads cost 0.1 times the uncached input rate. That means a cacheable Sol prefix costs $2.50 per one million tokens when written, then $0.20 per one million tokens when read from cache; a Luna prefix costs $0.125 per one million tokens when written, then $0.01 per one million tokens when read. A write followed by one full cached read can still be cheaper than processing the same prefix twice uncached, but a one-off cache write with no reuse is more expensive than ordinary uncached input for that prefix.
A simple Sol example shows the accounting problem. If an application sends a stable 100,000-token policy-and-tool prefix once and never reuses it, a cache write for that prefix is priced at the cache-write rate, not the cheaper cached-read rate. If the same exact rendered prefix is reused several times within the eligible cache window and compatible settings remain stable, later requests can benefit from the cached-input rate for the reused portion. The output, reasoning tokens, tool calls, and any new uncached prompt suffix are still separately billable. Prompt caching lowers repeated computation for matching prefixes; it does not make the whole request free, validate the prompt content, or authorize reuse of tenant-specific information across governance boundaries.
For developers building cost dashboards, “input price” should be split into at least three line items: fresh input, cache writes, and cached reads. Collapsing those into one blended number hides whether savings came from genuine prefix reuse, shorter prompts, a cheaper model, fewer retries, lower reasoning effort, or a change in the task. It also makes incident analysis harder when a schema change, tool-ordering change, verbosity change, or compaction event unexpectedly reduces cache reuse. OpenAI’s caching documentation states that cache misses can be caused by model, tools, tool ordering, structured output schema, reasoning effort, verbosity, context management, or earlier input changes, so the accounting system needs enough dimensions to identify drift.
The 272,000-input-token threshold changes the whole request, not just the excess
OpenAI’s Sol and Luna model pages document a long-context pricing rule for requests above 272,000 input tokens: input and cache rates are multiplied by 2x, and output rates are multiplied by 1.5x for the full request. The phrase “for the full request” is operationally important. Teams should not price only the tokens above 272,000 at a higher rate unless OpenAI’s documentation for their account and model explicitly says otherwise. Under the cited model pages, crossing the threshold changes the applicable rates for the entire request.
| Scenario | Sol effective listed rate under cited rule | Luna effective listed rate under cited rule | Decision rule |
|---|---|---|---|
| Request at or below 272,000 input tokens | $2 input, $0.20 cached input, $2.50 cache write, $10 output per 1M tokens | $0.10 input, $0.01 cached input, $0.125 cache write, $0.50 output per 1M tokens | Use the standard listed rates, then add output, tools, retries, mode, regional processing, and review cost. |
| Request above 272,000 input tokens | 2x input/cache rates and 1.5x output rate for the full request | 2x input/cache rates and 1.5x output rate for the full request | Re-evaluate whether retrieval, summarization, chunking, or a smaller stable prefix can preserve accuracy while avoiding unnecessary long-context pricing. |
The shared large context window does not mean every task should use it. The Sol, Luna, and Astra model references share a documented 1,050,000-token context window, a maximum of 922,000 input tokens, and up to 128,000 output tokens. Those limits describe capacity, not recommended prompt size, model quality, latency, or economic fit. Long contexts can be appropriate for document-heavy analysis, repository-scale coding tasks, or agent state that cannot be safely compressed, but they can also bury relevant evidence, increase review burden, and trigger higher pricing. A cost-conscious design retrieves the minimum necessary evidence, preserves stable reusable prefixes where correct, and records when the long-context threshold is crossed.
Regional processing and processing mode can also change effective task cost. The cited model pages state that regional processing adds a 10% premium where available for Sol and Luna, that EU data residency for Sol and Luna is available only with Standard processing, that Batch and Flex are priced at 50% of Standard rates, and that Fast mode is 2x applicable rates. These controls are not interchangeable business knobs. A regulated workflow may require a particular data-residency posture; an interactive coding assistant may need lower latency; a non-urgent evaluation batch may tolerate delayed processing. The cheaper processing mode is not acceptable if it violates data-region rules, user expectations, workspace policy, or contractual commitments.
Tool calls may add separate charges according to the model references, so teams should not infer task cost from token pricing alone. A workflow that uses web search, file search, code interpreter, hosted shell, computer use, MCP, or other supported tools may incur additional tool-related costs and operational review obligations. If a model produces a low-token answer after several expensive tool steps, the API token line can understate the real cost of the task. Conversely, a higher-token answer that avoids unnecessary tool calls may be cheaper in a particular workflow. Cost evaluation should be task-based, not price-card-based.
Capabilities and limits that matter before migration
Sol and Luna share several published model characteristics, but they are not the same model and should not be treated as drop-in substitutes without validation. OpenAI lists both as accepting text and image inputs and producing text output. The API identifiers are exactly gpt-6-sol and gpt-6-luna. Their knowledge cutoffs differ: Sol is listed with an April 20, 2026 knowledge cutoff, while Luna is listed with a May 18, 2026 cutoff. A later cutoff is not a quality ranking and does not remove the need for retrieval, dated sources, verification, and human review in workflows that depend on current facts.
| Capability or limit | GPT-6 Sol | GPT-6 Luna | Operational consequence |
|---|---|---|---|
| API model ID | gpt-6-sol |
gpt-6-luna |
Use the exact identifiers documented by OpenAI; do not invent aliases or assume provider-specific deployment names. |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 | Use retrieval and source verification for current or high-stakes facts; cutoff date is not a reliability guarantee. |
| Context window | 1,050,000 tokens | 1,050,000 tokens | Large shared context does not imply identical performance, latency, cost, or suitability. |
| Maximum input | 922,000 tokens | 922,000 tokens | Above 272,000 input tokens, documented long-context multipliers apply to the full request. |
| Maximum output | 128,000 tokens | 128,000 tokens | Reasoning tokens count against output/context limits even when not visible, so applications must handle incomplete responses. |
| Inputs and outputs | Text and image input; text output | Text and image input; text output | Do not assume speech, transcription, embeddings, or realtime behavior from these model pages. |
OpenAI’s reasoning documentation states that GPT-6 Sol and Luna default to medium reasoning effort. The Sol and Luna model references list supported reasoning efforts as none, low, medium, high, xhigh, and max. Effort and processing mode are separate controls: OpenAI states that GPT-5.6 and GPT-6 models support standard and pro reasoning modes, while pro mode performs more model work, increases token usage and latency, and bills those tokens at the selected model’s standard rates. Teams should not assume that a higher effort or pro mode is faster, cheaper, or safer; it is a configuration to evaluate against the task.
Reasoning tokens are billed as output tokens and count against output and context limits even though raw chain-of-thought is not exposed. This changes how “short answer” tasks should be budgeted. A request can consume output budget internally before producing much visible text, and OpenAI’s reasoning documentation warns that max_output_tokens can cause a response to end as incomplete, potentially before any visible output is produced. Production code should treat an incomplete response as a state requiring explicit handling, not as a valid empty answer, successful refusal, or safe default.
// Recommendation: record cost-relevant settings with every Sol/Luna request.
// This is a logging pattern, not an OpenAI endpoint contract.
{
"model": "gpt-6-sol",
"surface": "responses_api",
"reasoning_effort": "medium",
"processing_mode": "standard",
"regional_processing": "none_or_configured_region",
"input_tokens": 0,
"cached_input_tokens": 0,
"cache_write_tokens": 0,
"output_tokens": 0,
"reasoning_tokens_visible_to_usage_only": true,
"tool_calls": [],
"long_context_threshold_crossed": false,
"response_status": "complete_or_incomplete_or_error",
"human_approval_required_before_side_effect": true
}
OpenAI’s documentation also gives a model-specific warning for function calling. Sol and Luna support Chat Completions function calling only when reasoning_effort is none; OpenAI recommends the Responses API for built-in tools and normal function-calling workflows. That distinction is easy to miss during migration from older integrations. A team that copies a Chat Completions configuration with function calling and then raises reasoning effort may find that the assumed behavior is unsupported under the cited model pages. Validate endpoint, effort, mode, tools, and schema together rather than changing only the model ID.
The supported API surfaces listed for Sol and Luna include Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. The model pages list supported Responses tools including web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Tool availability in documentation does not mean every application should grant every tool. A coding workflow might allow apply-patch operations only in a sandboxed branch with review; a research workflow might allow web search but block external publication; a desktop-assistance workflow might require explicit human approval before any consequential computer-use action.
The cited model pages also list unsupported features: Sol and Luna do not support Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions endpoints. That unsupported list should be part of the migration checklist. If an existing application depends on the Assistants API, speech transcription, custom embeddings, live audio interaction, or fine-tuned variants, the presence of a lower list price for Luna or Sol does not make the model a direct replacement. The engineering work may require an endpoint redesign, a separate model for the unsupported capability, or a decision not to migrate that workflow.
This article explains how to configure ChatGPT Work and Codex starting defaults, including models, reasoning, Fast Mode, roles, local access, cloud access, and entitlement boundaries. The Configure ChatGPT Work and Codex Starting Defaults: Models, Reasoning, Fast Mode, Roles, Local Access, and Cloud Access article is a focused companion for ChatGPT Work and Codex Access because it best supports discussion of ChatGPT Work and Codex access because it focuses on the exact admin and entitlement settings that govern user access.
Why list price is not task cost
The cleanest way to price a Sol or Luna workload is to start with a unit of work, not a token table. A unit of work might be “review one pull request,” “summarize one contract packet,” “answer one internal policy question with citations,” “triage one support ticket,” or “generate one lesson plan draft.” For each unit, measure the full path: prompt construction, retrieval, cache write, cached reuse, tool calls, reasoning tokens, visible output, retries, validation, human review, and downstream remediation. A lower input-token price is valuable only if the workflow still meets accuracy, latency, permission, and review requirements.
- Measure the baseline. Record current model, endpoint, prompt version, average and tail input/output tokens, tool calls, retry rates, incomplete responses, review time, and failure categories.
- Run Sol and Luna on representative traffic. Include typical cases, edge cases, adversarial cases, tool-selection cases, tool-argument cases, and cases that require escalation or refusal.
- Normalize cache behavior. Separate fresh input, cache writes, and cached reads; do not compare a warm-cache candidate against a cold-cache incumbent unless that reflects production reality.
- Apply long-context and processing adjustments. Mark requests above 272,000 input tokens, Fast mode, Batch or Flex mode, Standard processing, regional premiums, and tool charges.
- Evaluate quality and approval burden. Count not only success rate but also reviewer corrections, unverifiable citations, tool misuse, misleading status claims, and escalation misses.
- Decide with rollback in mind. A migration rollback should restore model, prompt/cache policy, tool configuration, state handling, and previous validated behavior together.
A procurement dashboard that multiplies average input tokens by the Luna price and stops there will understate cost in several common cases. Output can dominate cost when the model produces long explanations, code diffs, structured reports, or multi-step plans. Reasoning tokens are billed as output tokens, even when the user cannot see them. Tool use can add separate charges and review obligations. Retries can double or triple the cost of difficult cases. Long-context multipliers can change the full-request rate. Fast mode can double applicable rates. Regional processing can add a 10% premium where available for Sol and Luna. None of those factors is visible in the headline “$0.10 input” or “$2 input” number.
List price also excludes organizational cost. A stronger model that produces fewer review escalations may be cheaper for a legal-technology team than a lower-priced model that requires extensive attorney correction, but that conclusion must be proven with local evidence and does not authorize unsupervised legal advice. A coding agent that generates a plausible patch cheaply can still be expensive if maintainers spend hours untangling an unsafe change. An education workflow that drafts materials quickly still requires teacher review for age appropriateness, accuracy, and local policy. A security workflow must not grant broader tool authority merely because the per-token route is cheaper.
Operational warning: Treat Sol and Luna’s lower prices as inputs to a controlled routing decision, not as permission to widen access, lower approval thresholds, or skip evaluation. Human approval remains mandatory for external messages, submissions, payments, purchases, bookings, destructive actions, permission changes, publication, legal commitments, campaign launches, code deployment, and other consequential operations.
Model selection should combine price, support matrix, and governance
OpenAI’s launch framing positions Astra as the strongest overall GPT-6 model, while Sol and Luna offer lower-cost alternatives with different capability and price tradeoffs. That hierarchy should guide, not replace, evaluation. Luna may be attractive for focused high-volume work because its listed input and output prices are far below Sol’s. Sol may be a better candidate for more complex coding, agentic workflows, or tasks where Luna’s lower cost does not meet the quality bar. Astra may remain appropriate for the hardest end-to-end work. The right route depends on measured performance under the application’s own prompts, tools, data, permissions, and review process.
The shared context and tool lists can tempt teams to build a single “GPT-6 configuration” and swap model IDs dynamically. That is risky. Sol and Luna list none as a supported reasoning effort, while Astra’s cited model page does not list none. Sol and Luna have different knowledge cutoffs from Astra. Pricing differs sharply. Availability differs across product surfaces. Unsupported endpoints remain unsupported. A robust routing layer validates each model’s allowed efforts, endpoint support, tool policy, regional processing requirements, and fallback behavior before sending traffic.
For Work and Codex users, OpenAI’s Help Center separates product access from API billing. Sol and Luna are models for ChatGPT Work and Codex and are not available in ordinary Chat conversations at the time of the cited Help Center update. Availability depends on plan, workspace settings, role permissions, and rollout access. The help article also states that Codex preserves a manually selected model, that desktop Work/Codex picker defaults are separate from ordinary Chat defaults, and that Codex CLI does not use the desktop slider. A workspace starting default does not grant access to a model unavailable to a user’s role.
That product-surface separation has practical consequences for enterprise administrators. Changing a Work or Codex starting model can affect user experience, but it does not rewrite API integrations, change API-key billing, or override role-specific controls. ChatGPT plan usage and API-key billing are separate. A team may test Luna in a desktop Work flow, Sol in Codex, and Astra through an API-backed service, but those tests do not share the same billing model, tool surface, permission model, or administrative controls. Administrators should document which surface was tested before applying conclusions elsewhere.
Security teams should review tool authority independently from model choice. The Sol and Luna references include powerful tools such as code interpreter, hosted shell, apply patch, computer use, MCP, and web/file search in the Responses tool ecosystem. Lower token prices can increase usage volume, which can increase the number of opportunities for prompt injection, unsafe tool arguments, accidental disclosure, or unauthorized side effects. Least privilege, sandboxing, egress controls, file permissions, audit logging, and human approval gates should remain in place even when the model shows stronger benchmark results or lower alignment-challenge failure rates in OpenAI’s evaluations.
Legal, compliance, health, finance, and youth-facing workflows need stricter thresholds than ordinary drafting tasks. A model can help organize information, produce a draft, summarize a policy, or identify questions for review, but it should not be positioned as providing personalized legal advice, medical judgment, financial direction, or unsupervised decisions about people. Current facts require retrieval and verification; regulated decisions require qualified human review; external communications require approval; and sensitive information should be minimized, redacted, or handled only under organization-approved controls. The cost advantage of Luna or Sol does not change those obligations.
A practical cost worksheet for Sol and Luna evaluations
The following worksheet is a recommendation for internal evaluation, not an OpenAI billing interface. Its purpose is to prevent teams from mistaking list price for workload cost. Use actual usage fields from the API response and official billing records where available, and keep separate rows for different model IDs, effort settings, processing modes, regional settings, tool policies, and prompt versions. If a request crosses 272,000 input tokens, apply the documented long-context multipliers for the full request under the cited model pages.
| Cost component | What to record | Why it changes the decision |
|---|---|---|
| Fresh input tokens | Uncached prompt tokens processed at the model’s input rate | Shows the baseline cost of prompt construction, retrieval stuffing, and conversation history. |
| Cache-write tokens | Tokens written to cache at 1.25x the uncached input rate | Prewarming or first-use writes can raise cost unless later reuse occurs. |
| Cached-input tokens | Reused prefix tokens billed at 0.1x the uncached input rate | Shows whether stable prefixes are actually reducing repeated computation. |
| Visible output tokens | User-visible response tokens | Long reports, code, JSON, and explanations can dominate task cost. |
| Reasoning tokens | Output-billed internal reasoning tokens shown through usage accounting where available | Higher effort can improve some tasks but increase cost and latency; raw chain-of-thought is not exposed. |
| Tool calls | Tool type, count, and any separate tool charges | Search, code execution, shell, computer use, and file tools can change both cost and risk. |
| Processing mode | Standard, pro reasoning mode, Batch, Flex, or Fast where applicable | Batch/Flex and Fast have documented pricing effects; pro mode increases model work and token usage. |
| Regional processing | Whether regional processing is used and whether a premium applies | Regional controls may be required for governance even when they add cost. |
| Retries and incomplete responses | Retry count, cause, and whether incomplete responses had usable visible output | A cheap first attempt can become expensive if difficult cases require repeated calls or manual rescue. |
| Human review time | Reviewer minutes, escalation category, and correction severity | Task cost includes the labor required to make the output safe, accurate, and usable. |
A disciplined team can use this worksheet to compare three routes: Luna for high-volume drafts, Sol for complex coding or agentic tasks, and Astra for the hardest cases that fail lower-cost routes. The comparison should include “no migration” as an option. If Luna’s lower token price produces more incorrect tool calls, missing citations, or review escalations, it may not be cheaper for that workflow. If Sol produces fewer retries than Luna at a higher unit price, it may win on total task cost. If Astra is required only for a small share of escalated cases, a tiered route may control cost without lowering the quality bar for difficult work.
Prompt caching can improve that tiered route only when the prefix is stable and appropriate to reuse. OpenAI’s caching announcement describes higher cache hit rates by default, a Prompt Caching Dashboard, diagnostics, explicit breakpoints, prewarming, and append-only instruction/tool patterns as controls for developers. Its documentation also states that, on GPT-6, reasoning effort can be changed through an appended configuration_update without rewriting the earlier cached prefix, while top-level reasoning changes can affect cache reuse. The safe pattern is to keep stable developer messages, tool definitions, schemas, and ordering unchanged when correct, then append new instructions later in the context rather than rewriting the prefix.
Cache optimization should never override semantic correctness. If an instruction is obsolete, unsafe, tenant-specific, or legally inappropriate for a later request, preserving it for a cache hit is the wrong optimization. If a tool definition should be removed for permission reasons, do not retain it merely to improve reuse; where OpenAI’s documentation recommends allowed_tools or tool_choice: none instead of removing definitions, apply that pattern only when it matches the application’s security model and does not leave unauthorized call paths available. Cache design is a FinOps tool, not a substitute for access control, tenant isolation, or content governance.
Migration teams should validate the support matrix before changing production traffic
A safe migration plan starts by listing every feature the existing workflow uses and checking it against the Sol and Luna model pages. If the current system uses Responses with structured outputs and file search, Sol or Luna may be candidates subject to evaluation. If it uses legacy Completions, Realtime, Live, Assistants, fine-tuning, embeddings, speech, or transcription, the cited Sol and Luna pages say those features are not supported. If it uses Chat Completions function calling with reasoning effort above none, the Sol/Luna function-calling limitation becomes a redesign issue, not a simple model replacement.
The support check should also include output handling. Sol and Luna can produce up to 128,000 output tokens according to the model pages, but that capacity should be bounded by product needs and safety rules. Set max_output_tokens high enough to avoid accidental empty incomplete responses in legitimate tasks, yet low enough to control runaway verbosity and review burden. When an incomplete response occurs, the application should preserve evidence, avoid external side effects, and either retry under a controlled policy, ask for clarification, escalate to a human, or fail safely. Do not silently publish, merge, purchase, message, or submit based on a partial response.
Finally, rollout evidence should be separated from OpenAI’s public benchmark claims. OpenAI reports capability improvements for professional work, factuality, coding, computer use, and collaboration style, and it provides evaluation-specific figures in the launch article. Those results are useful signals, but the launch page itself notes that evaluations may differ from production ChatGPT because of system prompts, tools, and deployment differences. Competitor comparisons and cost-per-task figures also have stated limitations. Treat vendor benchmarks as hypotheses to test locally, not as approval to skip shadow traffic, canary deployment, monitoring, rollback, and human review.
The same caution applies to safety and alignment reporting. OpenAI’s deployment-safety materials and Sol/Luna appendix describe results under defined evaluation conditions, including challenging situations designed to stress failure modes. Those tests are not typical-use incident-rate estimates and do not prove universal reliability, deception resistance, or safety in a customer’s environment. A model that performs better in a challenge set can still make consequential mistakes, misuse a tool, overstate completed work, or produce unverifiable claims. Production adoption still requires sandboxing, access controls, audit logging, incident response, and qualified human approval for high-impact actions.
Benchmarks, caching, and alignment: how to read the strongest claims without over-reading them
OpenAI’s launch article for GPT-6 Sol and GPT-6 Luna makes three categories of claims that should be handled separately in any technical review: benchmark performance, prompt-caching economics, and alignment or safety behavior. Benchmark figures describe results under specified evaluation conditions, not universal production outcomes. Caching claims describe reusable-prefix billing and control mechanisms, not semantic memory or guaranteed savings. Alignment evaluations stress deliberately difficult scenarios and should not be converted into typical-use failure rates.
The practical implication is that a buyer, developer, or enterprise administrator should not ask only whether Sol or Luna “won” a public score. The more useful question is whether the launch evidence points to a testable local hypothesis. For example, OpenAI’s AutomationBench and DeepSWE results may justify a coding-agent pilot, while the 90% cached-read discount may justify redesigning stable system prompts and tool schemas. Neither fact is enough to remove human review from code merges, legal submissions, customer communications, access-control changes, or high-impact business operations.
AutomationBench: a professional-work signal, not an automation guarantee
OpenAI reports GPT-6 Sol results on AutomationBench in the launch article, including Sol at xhigh effort with a 33.2% result and an attributed cost of $0.27 per task. That claim is useful because it combines a dataset, a reasoning-effort setting, and a cost-per-task figure, rather than presenting a raw score alone. It is also limited: the launch article itself notes that OpenAI evaluations may differ from production ChatGPT because of system prompts, tools, and other deployment differences.
AutomationBench should therefore be read as evidence that OpenAI sees Sol as stronger than earlier lower-cost tiers for certain professional automation tasks, not as evidence that one-third of a company’s internal workflows can be safely delegated. Real automation work includes identity boundaries, source-system permissions, stale data, partial failure recovery, audit logging, exception handling, approval steps, and downstream effects. A benchmark task can measure whether a model completes a defined challenge; it does not measure whether a production system should be authorized to send the email, approve the refund, push the change, or close the ticket without review.
A conservative evaluation rule is to preserve the benchmark’s structure when designing a pilot: record model ID, effort, tool availability, system prompt, task class, cost per completed task, retry rate, human correction time, and failure severity. If a local workflow requires external side effects, the first pilot should run in shadow mode or a sandbox. The output can be scored against historical human outcomes, but execution authority should remain with an authorized person until the system has passed representative, adversarial, and operational evaluations.
Agents’ Last Exam: challenge-set results are not operating-room reliability
OpenAI’s launch discussion also includes agent-oriented evaluations such as Agents’ Last Exam. The important editorial point is not only the number reported in a vendor launch post, but the nature of the test. Agentic evaluations are typically designed to stress planning, tool use, persistence, and recovery across multi-step tasks. Those properties are relevant for founders building software agents, security teams assessing tool authority, and enterprise administrators deciding where to require approvals.
However, challenge evaluations are intentionally difficult and distribution-specific. A higher score on an agent challenge does not mean the model has a corresponding probability of success on your help-desk automation, finance reconciliation workflow, litigation-support review queue, or classroom feedback assistant. It also does not establish that the model will avoid all consequential mistakes when connected to privileged tools. The right interpretation is comparative and conditional: under OpenAI’s stated test setup, a model performed in a certain way; your system must test whether the same behavior holds with your prompts, tools, files, retrieval, roles, and approval controls.
For production migration, teams should distinguish “task solved in evaluation” from “operation safe to execute.” A model may produce the right plan while choosing an unauthorized tool, using an outdated document, hallucinating a missing confirmation, or failing to escalate a high-risk exception. The benchmark score can support prioritization, but it cannot replace least-privilege tooling, allowlists, confirmation screens, audit logs, rollback procedures, and human approval for consequential operations.
Factuality claims: useful direction, but retrieval and verification still matter
OpenAI says the new models improve factuality in the launch coverage, and the Sol/Luna appendix in the GPT-6 Astra system-card materials covers hallucinations and related safety behavior. These claims should be treated as evaluation results under defined conditions. They do not mean the model’s knowledge is current, that a longer knowledge cutoff is a quality ranking, or that retrieval can be skipped for current law, finance, medicine, product policy, security advisories, education rules, or enterprise procedures.
The model-reference pages list different knowledge cutoffs for the GPT-6 family: Sol is April 20, 2026, Luna is May 18, 2026, and Astra is April 30, 2026. A later cutoff is not a general superiority claim and does not remove the need for dated sources. For example, a legal-technology workflow that summarizes a filing rule must still retrieve the current court rule and preserve citation evidence. A security workflow that explains an exploit trend must still consult current advisories and internal telemetry. A classroom or youth-safety workflow must still follow the institution’s current policy and escalation procedure.
A practical factuality test should include questions with known answers, questions requiring “I do not know,” questions with conflicting sources, and questions where the correct answer depends on the date. Automated grading can catch exact-answer failures, but human review is still needed for subtle citation misuse, overconfident caveats, jurisdictional ambiguity, and missing escalation. OpenAI’s evaluation claims can justify running that test on Sol and Luna; they should not be treated as a substitute for it.
FrontierCode and DeepSWE: coding benchmarks require mergeability checks
OpenAI reports coding and software-engineering results in the launch article, including DeepSWE 1.1 results where Sol at max effort reaches 68.8% and Luna at max effort reaches 66.6%. These figures are notable because they show Luna, despite its much lower list token price, being presented as competitive in a specific coding evaluation at a specified effort. They are still benchmark results, not a guarantee that Luna will be the right default for every repository, language, test suite, or developer workflow.
DeepSWE-style results are most useful when translated into a mergeability evaluation. A patch-generating model should be assessed on whether it identifies the actual defect, edits the right files, preserves public APIs, passes existing tests, adds appropriate tests when needed, avoids broad rewrites, and produces a reviewable diff. A benchmark score cannot tell you whether your monorepo’s build system, internal libraries, secret-handling rules, code owners, and deployment gates are compatible with an automated patch workflow.
FrontierCode and related coding benchmarks should also be separated from operational authority. A model that solves hard coding tasks may still propose unsafe migrations, mishandle credentials in logs, delete compatibility code, or miss licensing constraints. Security teams should treat generated patches as untrusted code until reviewed and tested. Enterprise administrators should not allow a cheaper or stronger coding route to bypass branch protections, CI policy, code-owner review, software-composition analysis, or change-management requirements.
The model-reference pages add a configuration caveat for developers choosing an API surface. OpenAI documents that Sol and Luna support Chat Completions function calling only when reasoning_effort is none, and OpenAI recommends the Responses API for built-in tools and normal function-calling workflows. That matters for coding agents because tool calling, patch application, shell access, file search, and hosted execution are not interchangeable with plain text completion. A migration that changes only the model ID can silently inherit invalid assumptions about effort, tools, or endpoint support.
OSWorld and computer use: environment fidelity is the hard part
OpenAI’s launch discussion includes computer-use claims such as OSWorld-style evaluation results. These are relevant because many Work, Codex, and API-backed systems increasingly ask models to operate across graphical interfaces, documents, browsers, terminals, and hosted tools. A computer-use score can indicate progress in perceiving screens, planning actions, and recovering from interface state, but it cannot establish that an enterprise desktop automation is ready for unattended operation.
The production gap is especially important for computer use because real environments contain account-specific permissions, pop-ups, rate limits, localization, disabled controls, stale sessions, multifactor authentication, untrusted web content, and irreversible actions. A model may perform well in a controlled environment and still click the wrong destructive control when a live application changes. For regulated or high-impact contexts, the safe default is to require explicit confirmation before submission, purchase, publication, booking, permission change, account update, data deletion, or external message delivery.
Teams evaluating Sol or Luna for computer use should log the environment, allowed tools, screen state, action sequence, stop condition, and human interventions. The evaluation should include decoy controls, disabled actions, ambiguous UI labels, and cases where the correct behavior is to ask for help. A lower-cost model can make high-volume computer-use pilots more affordable, but lower cost should never broaden tool authority by itself.
Collaboration style: helpfulness must be tested against your review process
OpenAI says Sol and Luna improve collaboration style, which matters for knowledge workers and educators as much as for developers. Collaboration style covers behaviors such as asking clarifying questions, structuring work, accepting corrections, explaining uncertainty, and adapting to user intent. Those are valuable in ChatGPT Work and Codex because users often want a thinking partner rather than a one-shot answer.
The caveat is that a pleasant collaboration style can mask weak evidence. A model that sounds cooperative may still fail to cite sources, skip an edge case, accept a flawed premise, or overfit to a user’s preference. For founders and enterprise teams, the evaluation should score not only tone but also decision hygiene: Does the model identify missing inputs? Does it distinguish facts from assumptions? Does it ask for approval before consequential steps? Does it preserve a review trail that another person can audit?
Educators and parents should apply a similar distinction. A model that encourages a student can still provide overconfident explanations or do too much of the work. The safer design is to prompt for tutoring, feedback, and source-checking rather than direct completion of graded assignments. For youth-facing contexts, institutions should retain age-appropriate supervision, content policy, and escalation procedures; a launch claim about style is not a child-safety certification.
This article covers OpenAI’s enterprise workspace updates including Model Test, Codex policy audit logs, Groups Admin API, and group manager controls for proving access and preserving audit evidence. The OpenAI Ships Model Test, Codex Policy Audit Logs, Groups Admin API, and Group Managers for Enterprise Workspaces article is a focused companion for Model Evaluation Design because among the candidates, this is the closest operational match for model evaluation design because it specifically includes OpenAI’s Model Test feature rather than generic data-science model evaluation prompts.
Why challenge evaluations are not production failure rates
The most common mistake in reading launch benchmarks is to treat an evaluation percentage as a production probability. A score of 68.8% on a software-engineering benchmark does not mean a model will successfully complete 68.8% of your backlog tickets. A lower rate of misleading claims in an alignment challenge does not mean the model will mislead users at that rate in ordinary use. The dataset, prompt, tools, scoring rules, effort level, and environment define what the number means.
OpenAI’s own source notes reinforce this boundary. The launch article states that OpenAI evaluations may differ from production ChatGPT because of system prompts, tools, and other deployment differences. The Deployment Safety Hub materials emphasize that results are evaluations under defined conditions and are not proof of universal behavior, zero risk, or production incident rates. Those caveats are not minor legal footnotes; they are operational requirements for anyone deploying the models in real systems.
A production system has failure modes that a benchmark may not measure: retrieval returning the wrong source, a tool schema changing, a role losing access, a cache miss increasing latency, a user pasting confidential data, a long-context multiplier changing cost, a response ending as incomplete, or a reviewer approving a polished but wrong answer. Because of those failure modes, public benchmarks should feed a local evaluation plan, not replace one.
| OpenAI-reported area | What it can support | What it does not establish | Local validation step |
|---|---|---|---|
| AutomationBench and agent evaluations | Prioritizing professional-work and agent pilots | Permission to run unattended business actions | Shadow-mode task replay with approval-gate scoring |
| DeepSWE and coding benchmarks | Testing Sol or Luna in coding-agent workflows | Mergeability, security, maintainability, or repository fit | CI, code-owner review, regression tests, and diff-quality scoring |
| Factuality evaluations | Comparing answer quality under controlled prompts | Currentness, legal validity, medical accuracy, or source authority | Retrieval-backed tests with dated citations and human review |
| OSWorld and computer use | Testing GUI and tool-use workflows | Safe live operation across real accounts and irreversible controls | Sandboxed action traces, decoy controls, and explicit confirmations |
| Alignment challenge tests | Comparing behavior under adversarial or stressful scenarios | Typical-use incident rates or guarantees against deception | Scenario testing, monitoring, escalation, and post-incident review |
Prompt caching: the 90% cached-read discount and the 1.25x write cost must be evaluated together
OpenAI’s prompt-caching announcement says GPT-6 prompt caching offers discounts of up to 90% on eligible cached input-token reads. The model-reference pages express this in model-specific prices: Luna’s published Standard API prices are $0.10 per million input tokens, $0.01 per million cached input tokens, and $0.125 per million cache-write tokens; Sol’s are $2 input, $0.20 cached input, and $2.50 cache write; Astra’s are $10 input, $1 cached input, and $12.50 cache write. Those figures reflect the general rule in the prompt-caching documentation for GPT-5.6 and later: cache writes cost 1.25 times the uncached input rate, while reads cost 0.1 times that rate.
The arithmetic is simple but often misapplied. A single cache write is more expensive than one ordinary uncached input pass because it is billed at 1.25x. If the same eligible prefix is then read once at 0.1x, the combined prefix cost is 1.35x instead of 2x for two uncached passes. That can be attractive for repeated stable prefixes, but it is not automatically cheaper for one-off prompts, frequently rewritten instructions, tenant-specific data that should not be reused, or workflows where output tokens, tool calls, retries, and long-context multipliers dominate the bill.
Illustrative prefix-only arithmetic, excluding output, tools, retries, processing mode, and long-context multipliers:
Uncached twice:
1.00x + 1.00x = 2.00x
Write once and read once:
1.25x + 0.10x = 1.35x
Write once and read five times:
1.25x + (5 * 0.10x) = 1.75x
Decision rule:
Caching becomes economically useful when the stable, eligible prefix is reused enough times to offset the 1.25x write cost and when preserving that prefix is semantically correct.
Developers should also account for the published long-context threshold. The model-reference pages state that requests above 272,000 input tokens trigger 2x input and cache rates and 1.5x output rates for the full request. That means a large retrieval bundle or conversation state can change the economics of both cached and uncached paths. A team comparing Sol, Luna, and Astra should calculate uncached input, cache writes, cached reads, visible output, reasoning tokens, tool charges, retries, processing mode, regional processing premiums where applicable, and the long-context multiplier separately.
This article explains GPT-6 Astra prompt caching, including explicit breakpoints, cache keys, a 30-minute TTL, long-context limits, and cost-control implications. The GPT-6 Astra Prompt Caching Guide: Explicit Breakpoints, Cache Keys, 30-Minute TTL, and Long-Context Cost Control article is a focused companion for Prompt Caching Cost Control because it directly matches prompt caching cost control and is especially relevant to the current article’s discussion of new caching controls and API economics.
What prompt caching reuses, and what it does not do
OpenAI’s prompt-caching documentation says caching is enabled by default for supported models and reuses an unchanged rendered prefix. The cache stores key-value tensors, not prompt tokens. It can include developer messages, tool definitions, conversation history, text, images, documents, and supported audio, depending on the request and model support. This is a compute-reuse mechanism, not a semantic memory layer.
That distinction matters for governance. A cache hit does not mean the model has validated the facts in the prefix, authorized the user to reuse tenant data, or determined that the instructions are safe. It also does not mean the whole prompt was cached. Actual reuse is measured through usage.input_tokens_details.cached_tokens, and a session does not guarantee a cache hit. Applications still need access control, data minimization, prompt-injection defenses, source verification, and approval flows.
For GPT-5.6 and later, OpenAI documents a minimum cacheable visible prefix of 1,024 tokens. Shared prefixes remain eligible for reuse within a 30-minute window, and the current prompt_cache_options.ttl value is 30m. The documentation describes this as a minimum eligibility period after the latest write or reuse, while OpenAI may retain entries longer. Teams should not describe the 30-minute setting as a guaranteed physical-retention ceiling, and privacy reviews should consider model, organization policy, regional processing, logging, and application storage separately.
OpenAI also notes that cache location is machine- and region-dependent and that caches are not shared across organizations or regional processing boundaries. The optional prompt_cache_key can support separate accounting or isolation, but it does not guarantee a cache hit or pin a request to a machine. Multi-tenant systems should not try to share prefixes across customers merely to improve cache economics; tenant isolation and billing transparency are more important than marginal reuse.
Designing prompts for cache reuse without corrupting semantics
OpenAI’s caching announcement emphasizes higher cache hit rates by default and additional controls such as explicit breakpoints, prewarming, append-only instruction and tool patterns, a Prompt Caching Dashboard, and diagnostics. The practical technique is to keep the stable prefix stable: model-invariant policy text, durable tool definitions, structured output schemas, and long-lived developer instructions should appear before volatile user turns, retrieved snippets, or task-specific changes.
However, preserving a prefix is only correct when the prefix remains true and authorized. If a policy changes, a tool definition becomes unsafe, a tenant’s permissions change, or a data classification changes, the prefix should be updated even if that causes a cache miss. FinOps teams should not pressure developers to keep stale instructions solely to maintain a cache hit rate. The cost of a wrong or unauthorized cached prefix can exceed the token savings.
The documentation identifies common cache-miss causes: model changes, tool definitions, tool ordering, structured-output schema changes, reasoning effort, verbosity, context management, compaction, service tier, cache key, and earlier input changes. Tool definitions, schemas, and ordering should remain stable when they are still correct. If a tool should not be callable for a particular request, OpenAI recommends using controls such as allowed_tools or tool_choice: none where applicable rather than removing definitions and changing the cacheable prefix.
GPT-6 adds an important reasoning-effort pattern. OpenAI states that reasoning effort can be changed through an appended configuration_update without rewriting the earlier cached prefix, while top-level reasoning changes can affect cache reuse. Sol and Luna default to medium reasoning effort, and their model pages list supported values including none, low, medium, high, xhigh, and max. Code should still validate model-specific support before sending a request, because effort options and endpoint behavior vary across the family.
Recommended cache-friendly pattern, expressed as a workflow rather than a complete API request:
1. Put stable developer policy, durable tool schemas, and output contract first.
2. Avoid reordering tools or rewriting schemas unless the behavior truly changes.
3. Append new task instructions later instead of editing the stable prefix.
4. For supported GPT-6 requests, append configuration_update to change reasoning effort.
5. Measure cached_tokens, total input tokens, output tokens, retries, and tool fees.
6. Treat lower cache hit rate as acceptable when policy, authorization, or correctness requires a prefix change.
Dashboard and diagnostics: useful evidence, not a truth oracle
OpenAI’s prompt-caching launch article refers to a Prompt Caching Dashboard and diagnostics tool that give developers more control over cache behavior. The diagnostics documentation says supported GPT-5.6-and-later models using the Responses API can compare a request against a recent completed response by setting prompt_cache_options.comparison_response_id. This asks for diagnostic metadata; it does not load the earlier conversation and does not change caching behavior.
The diagnostics feature returns types such as cache_hit, cache_miss, comparison_response_not_found, and unavailable. A diagnostic cache_hit means no comparison miss was detected; it does not mean the entire current prompt came from cache. A miss can report a reason and token estimates, with documented reasons including model, cache key, service tier, tools, text format, reasoning effort, verbosity, compaction, and input changes.
The caveats are operationally important. Diagnostics are best effort, return the first classified reason, do not block or fail the model request, and have no additional diagnostic fee. Any extra baseline or retry model requests remain billable. Records expire after a short period. OpenAI says the diagnostics feature is compatible with Zero Data Retention and stores configuration metadata, token-count estimates, and hashes rather than raw prompts or outputs for diagnostics, but application logs, tool logs, retrieval systems, and observability pipelines need their own privacy review.
A practical dashboard review should separate three numbers: cache-write tokens, cached-read tokens, and uncached input tokens. A rising cache-hit rate can still coincide with higher total cost if output grows, retries increase, tools add charges, Fast mode is used, regional processing premiums apply, or a request crosses the 272,000-input-token threshold. Conversely, a lower hit rate can be acceptable if the team intentionally changed a tool schema, policy prefix, or tenant isolation boundary.
| Signal | What it tells you | Common false interpretation | Recommended action |
|---|---|---|---|
cached_tokens |
How many input tokens were reused from cache | The entire request was cached or verified | Compare against total input tokens and task correctness |
cache_miss diagnostic |
The first classified reason diagnostics found | The only reason reuse failed | Check model, tools, schema, effort, verbosity, and earlier input changes |
| High cache-hit rate | Stable prefixes are being reused often | The system is safe, cheaper overall, or more accurate | Audit cost components, authorization, and output quality separately |
| Dashboard cost drop | Caching or prompt design may be reducing repeated input computation | The same savings will apply to all workloads | Validate by workload, tenant, mode, model, and traffic shape |
Customer caching examples should be treated as attributed case studies
OpenAI’s caching article includes customer-reported examples, including GitHub reporting more than a 50% reduction in prompt tokens requiring fresh processing against its previous baseline and a Manus example reporting a cache-hit-rate increase from roughly 85% to above 90%. These are useful case studies because they show the kinds of gains that can occur when stable prefixes, tool definitions, and request patterns are optimized. They should not be treated as promises for a different workload.
Two applications with the same model and price sheet can see very different outcomes. A coding assistant that reuses a large repository instruction prefix across many turns may benefit substantially. A customer-support bot that injects different retrieved records for every request may see less benefit. A legal review system may deliberately separate tenants, matters, and privilege groups in ways that reduce reuse but improve governance. A classroom tool may choose shorter prompts and more retrieval checks rather than large prewarmed prefixes.
Prewarming is similarly a cost and latency tactic, not a safety mechanism. It can prepare stable content for reuse, but it does not validate the content, authorize its use, or secure the data. Organizations should require approval before prewarming sensitive instructions, regulated material, or tenant-specific context, and they should avoid prewarming secrets, credentials, private keys, unapproved production data, or unnecessary personal information.
Alignment claims: lower rates in hard tests do not remove oversight
OpenAI’s launch article and the Sol/Luna appendix to the GPT-6 Astra system card discuss alignment evaluations, including improvements over GPT-5.6 counterparts in several tests and lower rates of misleading claims about coding work. The key caveat is explicit in the source findings: these alignment evaluations deliberately stress challenging situations and do not estimate typical-use failure rates. They also do not guarantee that an agent will never misreport progress, overstate completion, conceal uncertainty, or take an inappropriate action.
This distinction is especially important for Codex and agentic coding workflows. A lower rate of misleading coding-work claims in an evaluation is encouraging, but production systems should still require evidence of work: diffs, test results, command logs, failure traces, dependency changes, and reviewer-readable explanations. A model should not be accepted at its word when it says tests passed, a vulnerability was fixed, a migration is safe, or a deployment is complete. The system should verify through tools and human review.
The Deployment Safety Hub materials also emphasize remaining failures, evaluation awareness, monitorability limits, and the fact that absence of observed failures does not establish reliability across settings. That is a direct warning against treating safety benchmarks as permission for broad autonomy. Monitoring can help detect problems, but it cannot replace alignment, access control, sandboxing, approval gates, incident response, and rollback.
Operational recommendation: treat OpenAI’s alignment results as a reason to test Sol and Luna in controlled workflows, not as authorization to remove reviewers, expand tool permissions, or automate high-impact decisions. The approval burden should be set by action risk, data sensitivity, reversibility, and regulatory context, not by launch-day benchmark optimism.
What a responsible Sol/Luna evaluation should measure next
A serious evaluation should combine OpenAI’s reported evidence with application-owned tests. OpenAI’s evaluation-best-practices documentation recommends eval-driven development: define a task-specific objective, collect representative data, define metrics, run comparisons, and evaluate continuously. The evaluation should include typical cases, edge cases, adversarial cases, tool selection, tool arguments, and agent handoffs. Public benchmarks and vendor examples are inputs to that process, not deployment gates.
For a coding workflow, the test set should include routine bug fixes, ambiguous tickets, failing tests, dependency conflicts, security-sensitive code, and cases where no change should be made. For a research workflow, it should include source conflicts, outdated sources, missing evidence, and citation verification. For a business-process workflow, it should include permissions failures, partial data, approval-required cases, and user requests that should be refused or escalated. Each test should record model ID, effort, mode, tools, prompt version, cache state, cached tokens, output tokens, latency, retries, tool fees, human correction time, and severity of failure.
Teams should avoid building a new critical production dependency on OpenAI’s retiring Evals platform without a transition plan. OpenAI’s documentation states that the current Evals platform becomes read-only for existing users on October 31, 2026 and is scheduled to shut down on November 30, 2026. A safer approach is to maintain an application-owned evaluation harness, use supported current tooling where appropriate, and store enough metadata to reproduce decisions after model, prompt, or tool changes.
The most defensible launch response is therefore neither immediate migration nor blanket skepticism. Sol’s and Luna’s lower list prices, coding results, professional-work claims, and improved caching controls create credible reasons to run targeted pilots. Those pilots should be bounded by representative evaluation, staged rollout, monitoring, rollback, and human approval for consequential operations. That is the only way to convert launch claims into local evidence without confusing benchmark progress with production assurance.
Decision implications for individual users, developers, administrators, and enterprise teams
OpenAI’s launch of GPT-6 Sol and GPT-6 Luna creates a practical decision point rather than a simple upgrade instruction. Individuals using ChatGPT Work or Codex should first confirm that the model appears in the relevant Work or Codex surface for their plan, role, app, and rollout state; the launch and Help Center materials distinguish these surfaces from ordinary Chat, and OpenAI states that Sol and Luna are not yet available in ordinary Chat at launch. Developers should treat the API identifiers gpt-6-sol and gpt-6-luna as new candidates for measured routing, not drop-in replacements for every GPT-5.6, GPT-6 Astra, or earlier-model path. Administrators should separate workspace defaults, role eligibility, local permissions, browser or network controls, and Codex permissions, because OpenAI’s Help Center states that a configured starting default does not grant access to a model unavailable to a user’s role.
For knowledge workers, the immediate operational implication is to keep tasks inside the correct product boundary. If a user sees Luna in the desktop application, that does not imply that Luna is available in ordinary Chat, that Sol is available to the same user, or that API billing follows the user’s ChatGPT plan allowance. If a Codex user manually selects a model, OpenAI’s Help Center says Codex preserves that selection, while desktop Work and Codex defaults are separate from ordinary Chat defaults. A practical check is to record which app, workspace, role, and task surface produced the result before comparing quality, speed, or usage.
For developers, the most important implication is that the headline input and output prices are only two lines in the cost model. Sol’s listed Standard API prices are $2 per million input tokens and $10 per million output tokens; Luna’s are $0.10 per million input tokens and $0.50 per million output tokens. OpenAI describes both as 50% cheaper than GPT-5.6 promotional prices, but that statement does not guarantee lower total cost for a workflow that expands prompts, increases reasoning effort, emits longer outputs, retries more often, uses paid tools, writes prompt caches, runs in Fast mode, uses regional processing where available, or crosses the documented long-context threshold above 272,000 input tokens.
For enterprise security and platform teams, the launch should trigger a model-governance update rather than only a procurement update. Routing to a lower-cost model must preserve the same or stronger access controls, tool policies, data-region commitments, approval steps, audit logging, and incident response paths. Prompt caching can reduce repeated computation for matching prefixes, but OpenAI’s documentation is explicit that caching does not validate data quality, authorize data sharing, prevent prompt injection, or replace application controls. If a workflow handles confidential contracts, source code, personal data, security findings, financial decisions, health content, student records, or regulated operations, the model change should be reviewed as a system change with a rollback plan.
This article explains how to build an AI business value dashboard for ChatGPT Work and Codex using analytics on usage, spend, task mix, outcomes, and the Admin API. The Build an AI Business Value Dashboard with ChatGPT Work and Codex Analytics: Usage, Spend, Task Mix, Outcomes, and the Admin API article is a focused companion for AI Business Value Measurement because it is the strongest fit for measuring business value because it focuses on outcome-oriented analytics rather than generic enterprise AI adoption.
What each audience should do before adopting Sol or Luna
| Audience | Primary decision | Minimum check before use | Common mistake to avoid |
|---|---|---|---|
| Individual Work or Codex users | Whether Sol or Luna is appropriate for a specific task in the available product surface | Confirm model availability in the actual Work or Codex interface, task type, and workspace policy | Assuming desktop, Work, Codex, ordinary Chat, and API access are interchangeable |
| Developers | Whether gpt-6-sol or gpt-6-luna should be routed for an API workload |
Run representative evals with cost, latency, output quality, tool correctness, incomplete-response handling, and retry accounting | Replacing only the model ID while leaving incompatible effort, endpoint, cache, or tool assumptions in place |
| Workspace administrators | Which roles may use Sol or Luna in ChatGPT Work and Codex | Review starting model, reasoning level, speed, Fast Mode availability, new-chat behavior, role permissions, file access, and local/cloud permissions | Believing a workspace default overrides plan, rollout, or role restrictions |
| Security and compliance teams | Whether model routing or prompt caching changes risk boundaries | Validate data classification, tenant isolation, regional processing, audit logging, prompt-cache accounting, and human approval gates | Treating prompt-cache reuse as permission to share prefixes across users, tenants, organizations, or regions |
| Enterprise product owners | Whether Sol or Luna improves a business workflow enough to deploy | Compare against the current baseline using task-specific acceptance thresholds, reviewer workload, error severity, and total cost per successful task | Using public benchmark placement or token list price as the sole deployment gate |
Rollout checks before changing traffic, defaults, or user guidance
A safe rollout starts by inventorying where the model will be used. Teams should list every application surface, endpoint, system prompt, developer instruction, tool definition, schema, retrieval source, logging destination, and human approval point that participates in the workflow. OpenAI’s model pages state that Sol and Luna support Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching, while the cited pages do not list support for Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions. A migration plan that depends on an unsupported endpoint is not ready for production.
The next rollout check is reasoning configuration. OpenAI’s model references list Sol and Luna reasoning efforts as none, low, medium, high, xhigh, and max, with medium as the default. However, model-specific settings still need validation in code, because copying a configuration across the GPT-6 family can break: Astra does not list none effort in the cited model reference, and Sol/Luna Chat Completions function calling is documented as available only when reasoning_effort is none. For built-in tools and general function-calling workflows, OpenAI recommends the Responses API for reasoning models.
A third rollout check is long-context pricing exposure. The Sol and Luna model pages document a 1,050,000-token context window, a maximum of 922,000 input tokens, and up to 128,000 output tokens, but the shared window should not be treated as equal cost or equal fit. For requests above 272,000 input tokens, OpenAI documents 2x input and cache rates and 1.5x output rates for the full request. This means a summarization, legal-review, source-code, or document-analysis workflow can look inexpensive in a small sample and become materially different when full-length production prompts cross that threshold.
A fourth rollout check is incomplete-response handling. OpenAI’s reasoning documentation states that reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible. It also warns that max_output_tokens can end a response as incomplete, potentially before any visible output is produced. Production applications should handle incomplete as a controlled state with retry, escalation, or safe failure logic; they should not silently accept an empty or partial answer as a completed review, completed patch, completed filing, completed support response, or completed compliance analysis.
Recommended staged adoption path
- Baseline the current workflow. Record the current model, prompts, tool definitions, schemas, context size, reasoning settings, latency, token usage, cache behavior, reviewer effort, escalation rate, failure categories, and total cost per accepted task.
- Run offline representative evaluations. Use redacted, synthetic, public, or organization-approved data covering typical cases, edge cases, adversarial cases, tool-selection cases, tool-argument cases, and human-review cases.
- Shadow production inputs without side effects. Send copies of approved inputs to Sol or Luna while preventing external messages, submissions, purchases, bookings, account changes, deployments, destructive operations, or permission changes.
- Compare against explicit acceptance thresholds. Use task-specific rubrics for correctness, factual support, mergeability, tool-call validity, policy compliance, reviewer workload, latency, and total cost per successful task.
- Canary with least privilege. Route a small, reversible cohort only after offline and shadow evidence meets thresholds, and keep human approval for consequential actions.
- Expand gradually with monitoring. Increase traffic only if quality, safety, usage, cache, latency, and incident indicators stay within predefined limits.
- Rollback as a bundle when needed. Restore model selection, prompts, cache policy, tool availability, schemas, state handling, and prior validated behavior together; switching only the model ID can leave incompatible assumptions in place.
Representative evaluations should decide routing, not public benchmark rank
OpenAI’s launch article reports benchmark and evaluation results for Sol and Luna, including professional-work, coding, computer-use, factuality, collaboration-style, and alignment tests. Those values are useful evidence about the models under defined conditions, but OpenAI also states that its evaluations may differ from production ChatGPT because of system prompts, tools, and other deployment differences. The deployment decision for a legal research assistant, pull-request reviewer, procurement summarizer, student-feedback tool, customer-support triage system, or security-analysis copilot should be based on the organization’s task distribution and risk threshold.
A representative evaluation set should include successful ordinary cases and failure-seeking cases. For coding, that means not only whether a model writes plausible code, but whether the patch compiles, tests pass, dependency changes are justified, security-sensitive behavior is preserved, and the diff is reviewable. For research, it means whether claims are sourced, dates are handled correctly, uncertainty is expressed, and unsupported statements are flagged. For agentic workflows, it means whether the model chooses the right tool, passes valid arguments, respects tool authority, stops when permissions are missing, and asks for human approval before consequential action.
OpenAI’s evaluation best-practices documentation recommends eval-driven development: define a task-specific objective, collect representative data, define metrics, run comparisons, and evaluate continuously. It also notes that academic or generic benchmark scores alone are not substitutes for application-specific evals. Teams should treat pairwise comparison, classification, and criterion-based scoring as more operationally useful than vague open-ended ratings, while recognizing that LLM-as-judge systems can show position and verbosity bias. Human labels should calibrate automated graders before those graders influence rollout or routing.
Teams should also avoid building new critical workflows around the retiring OpenAI Evals platform without a transition plan. OpenAI’s evaluation documentation states that the current Evals platform becomes read-only for existing users on October 31, 2026 and is scheduled to shut down on November 30, 2026. A safer approach is to maintain an application-owned evaluation harness or use a currently supported alternative that records prompts, model IDs, settings, tool traces, token counts, cache state, grader versions, human decisions, and rollback evidence.
Evaluation evidence that should be captured for Sol/Luna decisions
| Evidence category | What to record | Why it matters |
|---|---|---|
| Model configuration | Model ID, endpoint, effort, mode, service tier, verbosity, prompt version, schema version, and tool list | Cache reuse, output quality, latency, and tool behavior can change when these settings change |
| Task quality | Pass/fail outcomes, rubric scores, reviewer notes, factual support, code-test results, and escalation decisions | Public benchmarks do not prove local fitness for a specific workflow |
| Tool behavior | Tool chosen, arguments passed, authorization checks, blocked calls, side-effect prevention, and error handling | Agentic success depends on correct tool use, not only final prose quality |
| Cost and usage | Uncached input, cache writes, cached reads, visible output, reasoning tokens, retries, tool fees, processing mode, and regional premiums | Total task cost can diverge from list input/output pricing |
| Safety and governance | Policy violations, sensitive-data exposure risks, prompt-injection outcomes, approval bypass attempts, audit-log completeness, and incident tickets | Model change cannot weaken compliance, security, or human-review requirements |
| Failure recovery | Incomplete responses, timeouts, malformed structured output, invalid tool calls, cache misses, and rollback success | Production readiness depends on controlled failure modes, not only successful examples |
Administrator controls: role permissions, starting defaults, and Work/Codex boundaries
Workspace administrators should treat Sol and Luna rollout as a permissions and governance exercise. OpenAI’s Help Center says availability depends on plan, workspace settings, role permissions, and rollout access. Owners and administrators can configure starting model, reasoning level, speed, Fast Mode availability, and new-chat behavior for Work and Codex, but those controls operate within the limits of plan and role eligibility. A starting default is not a license grant, not a promise of universal model visibility, and not evidence that every user can access the same model across desktop, web, mobile, Codex, and API workflows.
Administrators should document model access by role rather than by enthusiasm or seniority. For example, a legal-operations reviewer may need Sol or Astra for high-consequence contract analysis with strict human approval, while a high-volume internal summarization workflow may be evaluated on Luna if its outputs are nonbinding and reviewed. A developer working in Codex may need model access for code review but not permission to deploy, merge, change production secrets, alter identity policies, or run destructive commands. Model capability should not expand action authority.
Codex-specific settings deserve separate review. OpenAI’s Help Center notes that Codex preserves a manually selected model and that the desktop Work/Codex picker and defaults are separate from ordinary Chat defaults. It also states that Codex CLI does not use the desktop slider. Administrators should therefore avoid relying on a single visual default as proof of policy enforcement. A practical audit should compare workspace configuration, role permissions, local and cloud permissions, repository access, file access, network restrictions, browser controls, and the application logs that show which model and tool path were actually used.
Education and youth-adjacent environments should be especially conservative. The cited sources discuss plan and workspace availability, not a blanket suitability determination for minors, students, or sensitive educational records. Schools and education administrators should preserve existing student-data controls, parental or institutional policies, teacher review, accessibility checks, and escalation pathways. A model that can draft feedback or summarize material should not autonomously grade high-stakes work, make disciplinary decisions, disclose student records, or contact families without authorized human review.
Administrator rollout checklist
- Confirm eligible populations. Identify which users are on Plus, Pro, Business, Enterprise, or Edu plans for Work/Codex access, and separately verify any desktop Luna availability for Free or Go users where applicable.
- Separate Work, Codex, Chat, desktop, and API. Record the exact surface being governed; do not apply ordinary Chat assumptions to Work or Codex.
- Validate role-specific controls. Confirm that model defaults, reasoning levels, speed settings, Fast Mode availability, file access, local/cloud permissions, and browser/network rules align with each role’s business need.
- Preserve least privilege. Do not grant broader repository, file, browser, tool, billing, or administrative access merely because a lower-cost model is available.
- Require human approval for consequential outcomes. Keep approval gates for publication, legal commitments, payments, purchases, bookings, external messages, submissions, deployments, permission changes, and destructive actions.
- Record exceptions. Any user, team, or workflow that bypasses the default route should have a business justification, expiration date, owner, and audit trail.
Spend monitoring: cache-aware FinOps without assuming savings
Spend monitoring for Sol and Luna should start with the full request bill, not a headline token price. OpenAI’s model pages list Luna Standard pricing as $0.10 per million input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens, and $0.50 per million output tokens. Sol Standard pricing is $2 per million input tokens, $0.20 per million cached input tokens, $2.50 per million cache-write tokens, and $10 per million output tokens. Those figures should be tracked separately because cache writes, cached reads, uncached input, output, reasoning tokens, tool calls, retries, processing modes, regional processing, and long-context multipliers have different economics.
Prompt caching needs careful accounting. OpenAI’s caching documentation says GPT-5.6-and-later cache writes cost 1.25 times the uncached input rate and reads cost 0.1 times that rate. A write plus one full read can be cheaper than processing the same prefix uncached twice, but a single write without reuse is more expensive than ordinary uncached input for that prefix. Finance and engineering teams should monitor cache-write volume, cache-read volume, cache-hit-sensitive prefixes, diagnostic results, and the ratio of accepted outputs to total attempts. Savings should be measured per accepted task, not per optimistic prompt design.
The 30-minute prompt-cache window also requires precise interpretation. OpenAI documents prompt_cache_options.ttl as 30m for GPT-5.6 and later, and describes it as a minimum eligibility period after the latest write or reuse; OpenAI may retain entries longer. That means teams should not treat 30 minutes as a guaranteed physical-retention ceiling, a compliance deletion period, or a reason to include secrets in prompts. Cache location is machine- and region-dependent, caches are not shared across organizations or regional processing boundaries, and prompt_cache_key may help with accounting or isolation but does not guarantee a hit or pin a request to a machine.
Spend alerts should be tied to operational causes. A sudden cost increase might come from a model route changing from Luna to Sol, a reasoning effort change from medium to a higher effort, a service-tier change, Fast mode, long prompts crossing the 272,000-token threshold, a tool schema reorder that breaks cache reuse, a structured-output change, compaction of earlier turns, increased retries after incomplete responses, or a new workflow emitting long outputs. Cost dashboards that show only total dollars will not provide enough evidence for a safe rollback.
Recommended spend-monitoring fields
{
"workflow_id": "approved_internal_identifier",
"tenant_or_business_unit": "approved_accounting_bucket",
"surface": "api_or_work_or_codex",
"model_id": "gpt-6-sol_or_gpt-6-luna",
"endpoint": "responses_or_chat_completions_or_batch",
"reasoning_effort": "model_validated_value",
"reasoning_mode": "standard_or_pro_if_used_and_supported",
"input_tokens_uncached": 0,
"input_tokens_cached": 0,
"cache_write_tokens": 0,
"output_tokens_visible_and_reasoning_billed": 0,
"tool_calls": 0,
"retries": 0,
"processing_mode": "standard_or_batch_or_flex_or_fast_if_used",
"regional_processing": "none_or_regionally_configured",
"input_tokens_above_272k_threshold": false,
"response_status": "completed_or_incomplete_or_failed",
"human_review_required": true,
"human_review_outcome": "accepted_rejected_escalated_not_applicable",
"rollback_version": "model_prompt_tool_cache_state_bundle"
}
Human review, incident reporting, and rollback are still mandatory controls
OpenAI’s system-card materials and launch caveats should be read as a reminder that evaluation results do not establish zero risk, universal reliability, or production incident rates. The deployment safety materials emphasize remaining failures, evaluation awareness, monitorability limits, and the fact that the absence of observed failures does not establish reliability across settings. For enterprise teams, the practical rule is straightforward: model selection can improve a workflow, but it cannot remove approval gates for consequential outcomes.
Human review should be defined by action type, not by model confidence language. External messages, submissions, payments, purchases, bookings, destructive operations, permission changes, publication, code deployment, legal commitments, regulated decisions, account modifications, and high-impact recommendations require authorized human approval. A model may draft a response, propose a patch, summarize evidence, or identify issues, but the accountable person or system owner must verify the result before it affects another person, asset, system, legal position, or financial outcome.
Incident reporting should capture enough evidence to reproduce and triage the failure without exposing unnecessary confidential content. A report should include model ID, endpoint, effort, mode, prompt version, schema version, tool definitions, tool calls, cache settings, cache diagnostics if available, token usage, response status, user role, workspace policy state, surface used, and the human decision taken. Reports should avoid pasting secrets, credentials, private keys, personal identifiers, protected health information, student records, privileged legal material, or unrelated confidential documents. When sensitive content is essential to an investigation, teams should use approved secure evidence-handling procedures.
Rollback should be rehearsed before launch. Because cache behavior, prompt layout, tool definitions, schemas, service tiers, reasoning settings, and state handling can all affect output and cost, rollback must restore the validated bundle rather than only changing the model ID. If a Sol rollout changed prompts to preserve cache prefixes, introduced configuration_update messages, altered allowed tools, or modified structured-output schemas, reverting to the prior model without reverting those surrounding assumptions can produce new failures. A rollback runbook should specify who can trigger rollback, which metrics trigger automatic halt, which components are restored, how in-flight tasks are handled, and how reviewers are notified.
Operational triggers for pause or rollback
- Quality regression: A measured drop below acceptance thresholds for correctness, factual support, code-test pass rate, citation quality, mergeability, or reviewer acceptance.
- Safety regression: Unauthorized tool use, attempted approval bypass, sensitive-data exposure, prompt-injection success, misleading claims about completed work, or unreviewed consequential output.
- Cost anomaly: Unexpected increases in cache writes, uncached input, output/reasoning tokens, retries, long-context requests, tool charges, Fast mode usage, or regional-processing premiums.
- Availability mismatch: Users cannot access the configured model because of role, plan, workspace, surface, or rollout constraints.
- Incomplete-response spike: Increased
incompleteresponses, empty visible outputs, malformed structured outputs, or retries that create duplicated external actions. - Audit failure: Missing model ID, missing tool trace, missing reviewer decision, missing cache metadata, or unclear ownership for a high-impact action.
How to decide between adopting now, piloting, or waiting
Adopting Sol or Luna immediately can be reasonable for low-risk internal tasks where the organization can evaluate outputs, prevent side effects, and roll back quickly. Examples include internal draft generation, nonbinding summaries, exploratory code suggestions in a sandbox, and batch analysis of approved public or redacted material. Even in these cases, teams should record model ID, settings, cost, and reviewer acceptance so that early enthusiasm does not become unmeasured production dependence.
A controlled pilot is the better default for workflows with meaningful business impact. Customer-support triage, repository assistance, internal analytics summaries, knowledge-base drafting, procurement review, sales enablement, and education workflows often appear low risk until a hallucinated fact, unauthorized tool call, privacy exposure, or unreviewed publication reaches a user. Pilots should cap the user cohort, restrict tool permissions, define escalation rules, require human review for external outputs, and compare Sol or Luna against the current baseline using representative cases.
Waiting is appropriate when the workflow depends on unsupported endpoints, unvalidated tool behavior, strict regulatory interpretation, complex data residency requirements, or high-consequence autonomy. The cited model pages do not list Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions support for Sol and Luna. If a production system depends on one of those surfaces, the team should not force a migration by assuming support that is not in the cited model references. Waiting can also be the correct decision if the organization lacks evaluation data, logging, reviewer capacity, incident response, or rollback ownership.
The strongest adoption cases will usually share four traits: the task has a clear success metric, the input data is approved for the model route, the output is easy for a qualified person or deterministic system to verify, and the action authority remains limited. The weakest adoption cases are the opposite: ambiguous success criteria, sensitive or unapproved inputs, unverified factual claims, broad tool permissions, external side effects, and no accountable reviewer. Price reductions do not change that risk logic.
Conclusion: Sol and Luna expand the GPT-6 menu, but governance determines value
GPT-6 Sol and GPT-6 Luna give teams new cost-capability options below GPT-6 Astra, which OpenAI continues to position as its strongest overall model. The practical opportunity is real: lower list prices, Work and Codex availability for specified users, API identifiers for developers, and improved prompt-caching controls can make more workflows worth evaluating. The practical risk is also real: list prices do not guarantee workload savings, launch benchmarks do not prove local performance, alignment challenge tests do not remove oversight, and product-surface availability does not override plan, role, workspace, endpoint, or rollout limits.
The right response is disciplined experimentation. Individuals should verify the actual surface they are using. Developers should run representative evals before routing production traffic. Administrators should configure role permissions and defaults without confusing them with access grants. Enterprise teams should monitor spend with cache-aware accounting, preserve human approval for consequential actions, and rehearse rollback as a full configuration bundle. Sol and Luna may become valuable tools in many workflows, but their value will come from measured fit, controlled deployment, and accountable use rather than from the launch announcement alone.
Benchmark and alignment boundary: OpenAI reports results on Agents’ Last Exam and other named launch evaluations, but the reported challenge-set percentages are not an incident rate, not a production incident rate, and not a typical-use failure rate. They do not guarantee safety or reliability in a deployment.
Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!
Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.
Useful Links
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: Better prompt caching for GPT-6
- OpenAI Help Center: ChatGPT Work and Codex
- OpenAI API model reference: GPT-6 Sol
- OpenAI API model reference: GPT-6 Luna
- OpenAI Deployment Safety Hub: GPT-6 Astra system card and Sol/Luna appendix


