Choose GPT-6 Sol, Luna, or Astra for ChatGPT Work and Codex: Capability, Cost, Reasoning Effort, Permissions, and Review

Choose GPT-6 Sol, Luna, or Astra for ChatGPT Work and Codex: Capability, Cost, Reasoning Effort, Permissions, and Review
Choose GPT-6 Sol, Luna, or Astra for ChatGPT Work and Codex: Capability, Cost, Reasoning Effort, Permissions, and Review

Model choice is now a task, risk, and evidence decision

Choosing GPT-6 Luna, GPT-6 Sol, or GPT-6 Astra for ChatGPT Work, Codex, or an API-backed workflow should not start with the model name. It should start with the task, the cost of being wrong, the evidence you have for that task, and the controls around the work. OpenAI positions Luna as a lower-cost model for focused high-volume work, Sol as a higher-capability lower-cost option for complex coding and agentic workflows, and Astra as the strongest overall GPT-6 model for the hardest end-to-end work. Those descriptions are useful starting points, but they are not automatic routing rules for production systems, legal workflows, software deployment, regulated decisions, education settings, or enterprise knowledge work.

The practical decision is closer to a risk register than a model leaderboard. A short rewrite of an internal note, a codebase-wide refactor proposal, a customer-support draft, a compliance memo, a browser-using agent task, and a pull-request patch may all fit under “work,” but they have different failure modes. One task may tolerate a cheaper model if a human checks the final wording; another may require Astra plus strict tool limits, source citation, staged review, and rollback because a bad answer could affect security, contractual obligations, production code, or a person’s access to a service.

OpenAI’s September 22, 2026 launch article says GPT-6 Sol and GPT-6 Luna are faster and more affordable additions to the GPT-6 family, while GPT-6 Astra remains the strongest model overall. The same source reports improvements in professional work, factuality, coding, computer use, and collaboration style, but it also states that evaluations can differ from production ChatGPT because of system prompts, tools, and deployment differences. That caveat matters for every team making a model-selection decision: a public benchmark can justify testing a model, but it does not replace application-specific evaluation under your prompts, tools, permissions, data, latency constraints, and review process.

This guide treats the three models as options inside an operating system of controls. The right answer is not “use Luna for cheap tasks, Sol for coding, Astra for hard tasks.” The safer answer is: classify the task, define the unacceptable failures, choose the minimum model and reasoning settings that pass your representative evaluation, keep permissions independent of model choice, and escalate when evidence shows the model or configuration cannot meet the review threshold. For consequential work, the route must include authorized human approval even if the model produces an apparently polished answer.

For administrators, the first operational mistake to avoid is mixing up product surfaces. Ordinary Chat, ChatGPT Work, Codex desktop, Codex CLI, ChatGPT plan allowances, workspace defaults, role access, and API billing are related but separate concepts. OpenAI’s Help Center states that Sol and Luna are models for ChatGPT Work and Codex and are not available in ordinary Chat conversations at the time of the updated page. It also states that availability depends on plan, workspace settings, role permissions, and rollout access. A model appearing in one surface does not prove that it is available in every app, every workspace, every role, or every API project.

Separate Chat, Work, Codex desktop, Codex CLI, and API before comparing models

Before a team compares Luna, Sol, and Astra, it should identify where the work is actually running. OpenAI’s launch page says Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users; it also says Free and Go users can access Luna in the desktop application. At launch, OpenAI states that Sol and Luna are not yet available in Chat. That means an employee saying “I can use Luna” may be referring to a desktop Work or Codex surface, not ordinary Chat, and not necessarily an API deployment available to the company’s application.

Ordinary Chat should be treated as a separate conversational surface. The supplied OpenAI Help Center notes distinguish ordinary Chat defaults from Work and Codex defaults. If a user’s familiar chat experience does not expose Sol or Luna, that does not contradict Work or Codex availability. Conversely, a Work or Codex model picker does not imply that the same model is available for ordinary Chat messages, mobile workflows, browser sessions, or every plan. Teams that document procedures should name the surface explicitly, such as “ChatGPT Work desktop,” “Codex desktop,” “Codex CLI,” or “API Responses workflow,” instead of writing “use ChatGPT” as a catch-all instruction.

ChatGPT Work is a workspace-governed experience where owners and administrators can configure starting model, reasoning level, speed, Fast Mode availability, and new-chat behavior for Work and Codex, according to OpenAI’s Help Center. Those settings are important for reducing friction and nudging users toward approved defaults, but they do not override role access. OpenAI states that a workspace starting default does not grant access to a model unavailable to a user’s role. In practice, that means an admin default can be a starting preference, not a universal entitlement.

Codex desktop should also be separated from ordinary Chat because the Help Center says Codex preserves a manually selected model and that the desktop Work/Codex picker and defaults are separate from ordinary Chat defaults. This is relevant for coding teams that switch among quick analysis, patch generation, repository exploration, and local task planning. If a developer manually selects Sol for a complex coding session, the model selection may persist in Codex in a way that differs from the workspace starting default. Administrators should therefore audit real usage patterns rather than assuming that a configured default equals actual model use in every Codex task.

Codex CLI requires its own treatment. OpenAI’s Help Center states that Codex CLI does not use the desktop slider. If a company writes a model-selection policy around desktop UI controls, that policy can fail silently for CLI users. Engineering teams should document CLI configuration, allowed models, environment boundaries, repository permissions, review gates, and logging separately. A model that is appropriate for an interactive desktop coding assistant may not be appropriate for an automated CLI step unless the surrounding permissions, sandboxing, and human review are also appropriate.

The API is a separate economic and governance system. The model identifiers for Sol and Luna are `gpt-6-sol` and `gpt-6-luna`, according to OpenAI’s launch article and model references. API token pricing is not the same as ChatGPT plan usage. OpenAI’s Help Center specifically notes that ChatGPT plan usage and API-key billing are separate, and that Sol, Luna, and Astra can consume usage differently depending on task, input/output size, reasoning settings, and Fast Mode. A procurement team should not infer API cost from an employee’s plan allowance, and a developer should not infer ChatGPT availability from an API model page.

Surface or control What the OpenAI sources establish Decision rule for teams
Ordinary Chat OpenAI’s launch and Help Center notes state that Sol and Luna are not available in ordinary Chat at the cited launch/update time. Do not write procedures that assume ordinary Chat users can select Sol or Luna unless current workspace and product evidence confirms it.
ChatGPT Work Sol and Luna are available in Work for specified paid and organizational plans, subject to plan, role, workspace settings, rollout, and policy. Set defaults as guidance, then verify role-level access and review requirements for the tasks users actually perform.
Codex desktop Codex preserves a manually selected model, and the desktop Work/Codex picker is separate from ordinary Chat defaults. Audit actual selected models for coding workflows, not only workspace starting defaults.
Codex CLI OpenAI’s Help Center says Codex CLI does not use the desktop slider. Create separate CLI configuration, least-privilege, logging, and approval procedures.
API API billing uses model-specific token prices and settings; ChatGPT plan allowances are separate. Estimate cost from API usage details, reasoning tokens, tools, cache behavior, processing mode, retries, and long-context multipliers.

Define Luna, Sol, and Astra as positions, not automatic policies

OpenAI positions Luna as the most affordable of the three listed GPT-6 options in the supplied sources. The model reference lists Luna Standard API prices at $0.10 per million input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens, and $0.50 per million output tokens. The launch article describes Luna as part of a faster, more affordable Sol/Luna expansion below Astra. For high-volume focused work, Luna may be the first candidate to evaluate, especially where outputs are short, the review burden is manageable, and the task does not require the strongest available model.

That does not make Luna the default for every “simple” task. A short task can still be high risk if it involves legal commitments, regulated records, security access, a disciplinary decision, youth-safety concerns, medical information, financial instructions, or a public statement. A high-volume task can also become expensive if it triggers long outputs, repeated retries, billable reasoning tokens, tool calls, or long-context multipliers. Teams should evaluate Luna against the actual failure cost, not merely the apparent simplicity or the low input-token price.

OpenAI positions Sol as a stronger lower-cost option for complex coding and agentic workflows compared with Luna. The model reference lists Sol Standard API prices at $2 per million input tokens, $0.20 per million cached input tokens, $2.50 per million cache-write tokens, and $10 per million output tokens. OpenAI’s launch article reports benchmark examples including Sol xhigh at 33.2% on AutomationBench for $0.27 per task and Sol max at 68.8% on DeepSWE 1.1, while also warning that evaluations differ from production ChatGPT because of prompts, tools, and deployment conditions. These reported results can inform which model to test for coding and tool-using work, but they do not prove that Sol will be best for a particular repository, security policy, or agent architecture.

Sol is often the model to test when the task requires multi-step reasoning, repository navigation, code transformation, structured planning, tool coordination, or higher confidence than Luna provides in a representative evaluation. However, Sol should not automatically receive broader tool permissions. A coding assistant using Sol should still run with repository-level permissions, branch protections, review gates, test requirements, and explicit human approval before merge, deployment, destructive operations, credential changes, or production access. Model strength can reduce certain error rates in some settings, but it cannot replace access control.

OpenAI positions Astra as its strongest overall GPT-6 model. The model reference lists Astra Standard API prices at $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens, and $50 per million output tokens. Astra is therefore much more expensive per token than Luna and Sol under the listed Standard pricing, but a higher per-token price is not the same as a higher total task cost in every case. A model that solves a difficult task with fewer retries, fewer failed tool calls, or less human rework may be operationally preferable for selected work, while a lower-cost model may remain better for routine tasks that pass evaluation with acceptable review.

Astra is the candidate to evaluate for the hardest end-to-end work: tasks with ambiguous requirements, complex synthesis, high error cost, multiple tools, difficult code reasoning, demanding review standards, or executive, legal, security, and enterprise-administration implications. Even then, Astra should not be treated as authorization for unattended high-impact action. OpenAI’s Astra launch and system-card materials emphasize constraints around safeguards, monitoring, evaluation limits, and remaining failures. For consequential operations, the correct pattern is Astra plus controls, not Astra instead of controls.

Model OpenAI positioning in supplied sources Useful first evaluation candidates Operational warning
Luna Lower-cost GPT-6 option for focused high-volume work. Summarization, classification, routine drafting, structured extraction, first-pass coding assistance, and repetitive knowledge-work tasks with clear review. Low token price does not make a task low risk; do not use Luna for consequential decisions without human approval and task evidence.
Sol Higher-capability lower-cost option for complex coding and agentic workflows. Repository analysis, patch proposals, tool-coordinated work, complex debugging, multi-step planning, and agentic tasks under least privilege. Do not broaden file, browser, network, shell, or account permissions merely because Sol is stronger than Luna.
Astra OpenAI’s strongest overall GPT-6 model, intended for the most demanding work. Hard end-to-end tasks, high-stakes synthesis, difficult code or policy reasoning, and workflows where lower models fail representative evals. Higher capability does not remove the need for approvals, audit logs, sandboxing, rollback, source verification, and incident response.

Use shared context limits carefully: equal window size is not equal fit

The GPT-6 Luna, Sol, and Astra model references share a documented 1,050,000-token context window, a maximum of 922,000 input tokens, and up to 128,000 output tokens. They accept text and image inputs and produce text output. This shared scale can be valuable for large repositories, long policy documents, extensive research packets, or multi-file workflows, but it should not be misread as equal capability, equal speed, equal cost, or equal production fit. A million-token-capable model can still be the wrong choice if the task needs a smaller, cleaner, better-cited context and stricter review.

Large context creates governance and cost problems if teams use it as a substitute for retrieval design. Adding every available file, transcript, policy, ticket, and log line can increase exposure of confidential information, degrade relevance, increase review difficulty, and raise token cost. A more defensible approach is to retrieve the minimum approved material needed for the task, label sources, preserve access controls, and ask the model to separate cited facts from inference. For enterprise administrators, the context window is a capacity limit, not a permission model.

The model references also document long-context pricing multipliers. For requests above 272,000 input tokens, OpenAI documents 2x input and cache rates and 1.5x output rates for the full request. This matters when comparing Luna, Sol, and Astra for repository-scale or document-scale work. A model route that looks cheap at short context can become materially different once the entire request crosses the long-context threshold, especially if the workflow also uses reasoning, tools, retries, regional processing, or Fast mode.

Knowledge cutoffs are also not a simple ranking. The cited model pages list published knowledge cutoffs of April 20, 2026 for Sol, May 18, 2026 for Luna, and April 30, 2026 for Astra. A later cutoff does not mean the model is more capable overall, and it does not eliminate the need for retrieval, dated sources, and verification. For current legal, security, financial, product, scientific, or policy work, teams should use approved retrieval and cite current sources rather than relying on a model’s internal knowledge.

Attribute from cited model pages GPT-6 Luna GPT-6 Sol GPT-6 Astra Decision implication
Context window 1,050,000 tokens 1,050,000 tokens 1,050,000 tokens Same limit does not imply same quality, latency, price, or task fit.
Maximum input 922,000 tokens 922,000 tokens 922,000 tokens Large inputs require data minimization, access control, and cost forecasting.
Maximum output 128,000 tokens 128,000 tokens 128,000 tokens Long outputs can include billable visible text and billable invisible reasoning tokens where applicable.
Published knowledge cutoff May 18, 2026 April 20, 2026 April 30, 2026 Cutoff is not a freshness guarantee; current claims require retrieval and source verification.
Long-context multiplier Above 272,000 input tokens, 2x input and cache rates and 1.5x output rates apply to the full request. Cost comparisons must model the actual prompt size, not only the base token prices.

Account for token price, reasoning, tools, processing mode, and review cost

OpenAI’s published Standard API prices show the most obvious economic difference among the three models: Luna is listed at $0.10 input and $0.50 output per million tokens, Sol at $2 input and $10 output, and Astra at $10 input and $50 output. Cached input and cache-write prices also differ by model. For Luna, the cited prices are $0.01 cached input and $0.125 cache write per million tokens. For Sol, they are $0.20 cached input and $2.50 cache write. For Astra, they are $1 cached input and $12.50 cache write. These prices are token prices, not total task prices.

Total cost depends on the shape of the workflow. A task that produces long reasoning, long visible output, repeated retries, tool calls, or large retrieved contexts can cost more than expected even on a lower-priced model. A task that uses a higher-priced model but succeeds with fewer retries may be cheaper in human review time or incident risk. Finance teams should therefore compare total cost per accepted task, not only cost per million input tokens. Accepted-task cost should include uncached input, cache writes, cached reads, output tokens, reasoning tokens, tool charges where applicable, processing mode, regional processing premium where used, retries, failures, and human review time.

OpenAI’s reasoning documentation states that GPT-6 Sol and Luna default to medium reasoning effort. The model pages list Luna and Sol as supporting `none`, `low`, `medium`, `high`, `xhigh`, and `max`; Astra supports `low`, `medium`, `high`, `xhigh`, and `max`, and does not list `none`. This difference is important for routing and fallback. A configuration copied from Luna to Astra with `reasoning_effort: none` may not be valid for Astra. Production code should validate model-specific settings instead of assuming that every GPT-6 model accepts the same effort values.

Reasoning mode and reasoning effort are separate concepts in OpenAI’s reasoning documentation. GPT-5.6 and GPT-6 models support `standard` and `pro` reasoning modes, while supported effort values are model-dependent. Pro mode performs more model work, increases token usage and latency, and bills those tokens at the selected model’s standard rates. It should therefore be treated as an escalation setting, not a free quality switch. Teams should test whether pro mode improves acceptance rates enough to justify the extra token usage, latency, and review complexity for the specific task.

Reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible, according to OpenAI’s reasoning guide. This is a common source of underestimation in cost models and failure handling. If `max_output_tokens` is too low, a response can end as `incomplete`, potentially before any visible output is produced. Applications must handle an incomplete status explicitly by failing safe, retrying under a controlled policy, escalating to a human, or asking for a shorter approved output. They should not assume that every model response contains usable text.

Tool use further changes both cost and risk. The cited model pages state that Luna, Sol, and Astra support Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. They do not support Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions endpoints on the cited pages. The model pages also state that Sol and Luna support Chat Completions function calling only when `reasoning_effort` is `none`; OpenAI recommends the Responses API for built-in tools and normal function-calling workflows. Teams building new tool workflows should prefer the Responses API unless they have a documented reason and have verified model-specific support.

{
  "decision_record_template": {
    "task": "Describe the workflow in operational terms, not only the department name.",
    "surface": "Ordinary Chat, ChatGPT Work, Codex desktop, Codex CLI, or API Responses workflow.",
    "candidate_models": ["gpt-6-luna", "gpt-6-sol", "gpt-6-astra"],
    "reasoning_settings_to_test": "Model-specific effort and mode values only; do not assume one configuration works across all models.",
    "cost_components": [
      "uncached_input_tokens",
      "cache_write_tokens",
      "cached_input_tokens",
      "visible_output_tokens",
      "reasoning_output_tokens",
      "tool_charges_if_applicable",
      "long_context_multiplier_if_above_272000_input_tokens",
      "processing_mode",
      "regional_processing_premium_where_used",
      "retries",
      "human_review_time"
    ],
    "approval_rule": "Human approval required before external messages, publication, code merge, deployment, payments, permission changes, destructive actions, or regulated use.",
    "evidence_required": "Representative local evaluation with pass/fail criteria, failure taxonomy, and rollback plan."
  }
}

Classify work by error cost and authority before assigning a model

The most reliable model-selection framework begins by classifying the work, not by ranking the models. Each task should receive an error-cost rating, an authority rating, a data-sensitivity rating, and a review-burden rating. Error cost asks what happens if the answer is wrong. Authority asks what the model or agent can do through tools, files, browser access, shell access, account access, or publication workflows. Data sensitivity asks what information enters the context. Review burden asks how difficult it is for a qualified person to detect an error before the output causes harm.

A low-error-cost, low-authority task might be a first draft of an internal meeting summary using approved notes, with no external sending and a human editor. Luna may be a sensible first candidate to evaluate. A medium-error-cost task might be an engineering design review, a customer-facing draft, or a code-change plan that does not directly modify production. Sol may be a better first candidate if Luna fails to handle dependencies, edge cases, or tool instructions. A high-error-cost task might involve production code, legal interpretation, security configuration, financial commitments, regulated data, employee decisions, or public publication. Astra may be a candidate, but the model choice remains secondary to review, permissions, and evidence.

Authority must be separated from capability. A model may produce better plans when given tools, but every additional tool expands the possible blast radius. Browser/network controls, file permissions, role-specific controls, and local/cloud permissions remain separate from the Work/Codex model choice according to OpenAI’s Help Center. A lower-cost route never authorizes broader action scope, and a stronger model never justifies bypassing review. For example, a Sol-powered coding workflow should still be prevented from pushing directly to protected branches unless the organization’s normal change-management process explicitly allows that operation after human approval.

Data sensitivity can also invert the apparent model choice. A task involving confidential source code, personal records, privileged legal material, health information, student records, or security logs may require stricter workspace governance, access controls, redaction, regional processing review, or legal approval before model selection matters. Teams should not paste uncontrolled sensitive data into evaluation corpora or prompt examples. Use redacted, synthetic, public, or organization-approved inputs for testing whenever possible, and apply data-region and retention policies before enabling new workflows.

Review burden is often underestimated. Some model outputs are easy to check, such as a JSON transformation with a deterministic schema and known expected values. Others are difficult, such as a confident legal memo, a security root-cause analysis, or a proposed multi-file refactor. If reviewers cannot reliably detect errors, the task should be routed to stronger models, narrower tools, smaller evidence packets, or qualified expert review. In some cases, the correct decision is not to automate the task at all, but to use the model only for outlining questions, organizing evidence, or drafting non-final material.

Classification dimension Low-risk example Higher-risk example Model-selection consequence
Error cost Internal brainstorm alternatives for a non-sensitive document. Public legal, financial, health, safety, security, or employment-impacting decision support. Escalate model, evidence, and review as error cost rises; do not let model choice replace qualified judgment.
Tool authority Read-only file search over approved documents. Shell, browser, account administration, code modification, publication, or payment capability. Keep least privilege and require human approval for consequential actions regardless of model.
Data sensitivity Public documentation or synthetic test cases. Personal records, credentials, privileged files, regulated records, proprietary code, or security logs. Apply redaction, access control, data-region review, and logging policy before model evaluation.
Review burden Output can be checked by deterministic tests or a short rubric. Output requires expert judgment, source verification, or multi-step reproduction. Use stronger models and expert review, or narrow the task until review is feasible.

Use local evaluation as the gate, not public benchmarks or launch examples

OpenAI’s evaluation best-practices documentation recommends eval-driven development: define a task-specific objective, collect representative data, define metrics, run comparisons, and evaluate continuously. That guidance is especially important for Luna, Sol, and Astra because the public launch examples are not production guarantees. A benchmark result can indicate that a model is worth testing, but it cannot answer whether the model follows your internal policy, handles your source documents, uses your tools correctly, respects your review gates, or fails safely in your UI.

A useful evaluation set should include typical cases, edge cases, adversarial cases, tool-selection cases, tool-argument cases, and handoff cases. For Work and Codex, that may include routine knowledge-work prompts, ambiguous instructions, conflicting documents, incomplete repository context, failing tests, missing permissions, stale sources, and attempts to induce unauthorized actions. For legal-technology or education use, the evaluation should include source-verification tasks, jurisdiction or curriculum boundaries where relevant, and explicit uncertainty handling. For parents and youth-safety settings, the evaluation should avoid collecting unnecessary personal information and should route crisis or safety issues toward qualified real-world support rather than relying on an automated answer.

OpenAI’s evaluation documentation also warns that academic or generic benchmark scores alone are not substitutes for application-specific evals, and that automated scoring should be calibrated against human review. Pairwise comparison, classification, and criterion-based scoring often work better than vague open-ended judging. If a team uses an LLM judge, it should account for position and verbosity bias, use explicit rubrics, control response lengths where relevant, and validate judge agreement against human labels. The evaluation should produce a decision record, not merely a leaderboard screenshot.

OpenAI also notes that its current Evals platform is being deprecated: it becomes read-only for existing users on October 31, 2026 and is scheduled to shut down on November 30, 2026. Teams should not build a new critical production dependency on that retiring platform without a migration plan. A safer approach is to maintain an application-owned evaluation harness, exportable datasets, versioned prompts, model-configuration logs, reproducible scoring scripts, and human-review records that can move between supported tooling.

A practical model-selection gate can use three outcomes. First, “approved for default” means the model passes representative cases at the required quality, cost, latency, and review thresholds. Second, “approved for escalation” means the model is useful when a lower model fails or when the task is higher risk, but it is not the default because of cost or latency. Third, “not approved” means the model fails required thresholds or introduces unacceptable operational risk. This structure prevents teams from treating Luna, Sol, and Astra as a single ladder where every task automatically climbs upward or downward by price.

Recommendation: Treat every public benchmark, launch claim, and customer example as a hypothesis for local testing. Do not migrate a Work, Codex, or API workflow until representative evaluation shows that the selected model, reasoning settings, tools, permissions, and review gates meet the task’s acceptance criteria.

Set conservative escalation rules from Luna to Sol to Astra

Escalation should be explicit, observable, and reversible. Luna can be the first evaluated candidate for focused high-volume work where the task has clear instructions, limited authority, approved context, measurable output, and manageable review. Escalate from Luna to Sol when the evaluation shows recurring failures in multi-step reasoning, code understanding, tool selection, complex source synthesis, or instruction following that Sol resolves under the same controls. Escalate from Sol to Astra when the task remains too difficult, the error cost is high, or the review burden demands OpenAI’s strongest overall model and your evaluation supports the additional cost and latency.

Do not use cheapest-token routing as the primary rule. A routing system that always starts with Luna and retries on failure can be more expensive or risky than starting with Sol or Astra for known hard cases. Repeated failures can leak time, create inconsistent state, increase tool calls, frustrate reviewers, and generate confusing audit trails. For tasks with high authority or high error cost, the initial route should be chosen from the risk classification and evidence, not from the lowest input-token price.

Do not use model-name routing as the primary rule either. “Coding equals Sol” is too crude for real Codex work. A small unit-test explanation may be appropriate for Luna if it passes evaluation and has no write authority. A repository-wide migration plan may require Sol or Astra. A production patch should require tests, review, and approval regardless of model. Similarly, “executive work equals Astra” can waste budget if the actual task is a simple formatting pass over approved text, but Astra may be justified for high-stakes synthesis where lower models fail source fidelity or reasoning criteria.

Escalation criteria should be tied to failure categories. Examples include unsupported factual claims, missing citations, incorrect tool choice, unsafe tool arguments, incomplete response, schema violation, failure to ask for missing information, failure to preserve access boundaries, misleading claims about completed work, unreviewable reasoning, or unacceptable latency. OpenAI’s alignment evaluations for Sol and Luna include challenging situations and report improvements in some tests, including lower rates of misleading claims about coding work compared with GPT-5.6 counterparts, but those challenge tests are not typical-use incident rates or guarantees. Teams should still log and review whether an agent accurately reports what it did, what it could not do, and what evidence supports its output.

Rollback must restore more than the model identifier. A safe rollback restores model selection, prompt and cache policy, tool availability, state-handling behavior, reasoning settings, review gates, and prior validated behavior together. Switching only from Astra back to Sol, or Sol back to Luna, can leave incompatible assumptions in place if the prompt, tools, effort values, or output schema were tuned for the higher model. Every approved route should have a documented rollback bundle and an owner who can activate it when failure rates, cost anomalies, or safety incidents exceed thresholds.

Route decision Use only when evidence shows Do not use when Required control
Luna default Focused, high-volume work passes local evals with acceptable review cost and low authority. The task has high error cost, difficult review, broad tool authority, or repeated failures requiring retries. Human review for publication, external sending, regulated use, destructive action, or commitments.
Sol default or escalation Complex coding, tool coordination, or multi-step work benefits from Sol and remains within approved permissions. The workflow grants unnecessary authority or uses Sol’s capability as a substitute for tests and review. Least privilege, tool logs, test evidence, and approval before merge, deployment, or irreversible operations.
Astra escalation The task is among the hardest end-to-end workflows or lower models fail representative acceptance thresholds. The only justification is prestige, a single benchmark, or avoiding process controls. Expert review, audit logging, source verification, rollback, and explicit approval for consequential action.

Early decision checklist for Work, Codex, and API owners

The fastest safe way to begin is to create a one-page decision checklist for each workflow. The checklist should identify the surface, the candidate models, the task class, the authority level, the data class, the reasoning settings to test, the evaluation dataset, the review owner, the approval gate, the budget owner, and the rollback owner. This prevents a common failure pattern: one team changes the model picker, another changes the prompt, a third changes tool permissions, and no one can explain which change caused the observed quality, cost, or safety difference.

For ChatGPT Work administrators, the checklist should record workspace starting defaults, role-specific access, reasoning level, speed settings, Fast Mode availability, and new-chat behavior where these controls are available. Because OpenAI’s Help Center states that a starting default does not grant access to a model unavailable to a user’s role, the checklist should include a role verification step. It should also state which tasks require users to switch models manually and which tasks are prohibited from being handled in that surface because of data sensitivity, review burden, or missing controls.

For Codex desktop owners, the checklist should include the manually selected model behavior described in the Help Center. Developers need to know whether a repository task is expected to use Luna, Sol, or Astra, what reasoning settings are allowed, whether the model can edit files, whether shell or browser tools are available, and what must happen before a patch is merged. If Codex preserves a manually selected model, the team should avoid relying solely on starting defaults and should document how users confirm the selected model before high-risk work.

For Codex CLI owners, the checklist must not refer only to the desktop slider because OpenAI’s Help Center says the CLI does not use it. CLI workflows should specify model configuration, environment boundaries, repository scope, network access, credential handling, logging, and approval gates. They should never ask users to paste tokens, private keys, passwords, or uncontrolled secrets into prompts. If a CLI workflow can modify files or run commands, it should operate with least privilege and require human approval before destructive changes, permission changes, publication, deployment, or external communication.

For API owners, the checklist should include model IDs, endpoint support, pricing assumptions, reasoning settings, processing mode, cache policy, long-context threshold handling, tool charges, retry policy, incomplete-response handling, and audit logging. It should also include validation that the endpoint is supported for the chosen model. The cited model pages state that Luna, Sol, and Astra support Responses and Chat Completions, but not Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions endpoints. Building against an unsupported endpoint is not a model-selection problem; it is an architecture error.

  1. Name the surface: ordinary Chat, ChatGPT Work, Codex desktop, Codex CLI, or API. Do not use “ChatGPT” as a vague operational label.
  2. Classify the task: record complexity, error cost, tool authority, data sensitivity, review burden, and required evidence.
  3. Choose candidates: test Luna, Sol, and Astra only where the model is available and appropriate for the surface, role, and endpoint.
  4. Validate settings: confirm model-specific reasoning efforts, mode support, tool support, and endpoint support before running evaluations.
  5. Measure total cost: include uncached input, cached input, cache writes, output, reasoning tokens, tools, processing mode, regional premium where used, long-context multipliers, retries, and review time.
  6. Require evidence: use representative local evaluations rather than adopting public benchmark winners or launch examples as deployment gates.
  7. Keep permissions independent: never expand file, browser, shell, account, or publication authority merely because a stronger model is selected.
  8. Define approval gates: require authorized human approval before external messages, publication, code merge, deployment, payments, purchases, bookings, legal commitments, permission changes, destructive actions, or regulated decisions.
  9. Prepare rollback: restore model, prompts, reasoning settings, cache policy, tools, state handling, and review gates together.

This opening framework deliberately slows down model choice. That is the point. Luna, Sol, and Astra give teams a broader GPT-6 cost-capability range across Work, Codex, and API workflows, but broader choice also increases the chance of accidental policy drift. The defensible path is to treat model selection as an evidence-backed operating decision: map the surface, constrain authority, evaluate locally, account for total cost, and keep human review attached to consequential work.

Build the model decision matrix from risk, authority, evidence, and reversibility

Choose GPT-6 Sol, Luna, or Astra for ChatGPT Work and Codex: Capability, Cost, Reasoning Effort, Permissions, and Review — first editorial explainer visual

A practical Sol, Luna, and Astra routing policy should begin with the work being performed, not with a favorite model name or the lowest listed input-token price. OpenAI positions GPT-6 Luna for focused high-volume work, GPT-6 Sol for complex coding and agentic workflows, and GPT-6 Astra as the strongest overall GPT-6 model for the hardest end-to-end work. Those descriptions are useful starting hypotheses, but they are not permission to route regulated decisions, destructive code changes, customer commitments, or sensitive data through a cheaper path without controls.

The matrix below separates nine dimensions that teams usually collapse into one vague question: “Which model is best?” A safe answer depends on task complexity, error cost, latency pressure, volume, tool authority, data sensitivity, long-context use, review burden, and reversibility. A task may be simple but high-risk, high-volume but tool-authorized, or low-cost per token but expensive to review. The model route should reflect the highest-risk dimension, not the average of all dimensions.

Decision dimension Prefer Luna when… Prefer Sol when… Prefer Astra when… Required control regardless of model
Task complexity The task is narrow, repeatable, well-specified, and easy to verify, such as tagging support topics, extracting fields from approved documents, or drafting a first-pass summary. The task involves multi-step reasoning, code analysis, tool selection, repository navigation, or agentic planning with bounded authority. The task combines ambiguous goals, cross-domain reasoning, long dependency chains, high-value synthesis, or hard end-to-end problem solving. Use task-specific evaluation data rather than generic benchmark rankings.
Error cost Errors are inexpensive, visible, and easily corrected before any customer, legal, financial, security, or production consequence. Errors may waste engineering time or produce incorrect implementation guidance, but an expert reviewer can catch them before deployment. Errors could materially affect customers, contracts, security posture, regulated analysis, executive decisions, or production reliability. Human approval is mandatory for consequential actions and external commitments.
Latency pressure Fast, low-cost responses are important and quality requirements are modest enough for the task class. Moderate latency is acceptable for better reasoning on coding, tools, and multi-step workflows. Latency can be traded for stronger performance on the hardest work, subject to local measurement. Measure latency under the same effort, mode, tools, context size, and region used in production.
Volume The task runs at high volume and the organization has verified acceptable quality, escalation rate, and review cost. The task has meaningful volume but failures are costly enough to justify higher per-token spend for better capability. The task is low-volume, high-value, or high-risk enough that token cost is secondary to accuracy and reviewability. Track uncached input, cached reads, cache writes, output and reasoning tokens, retries, tool charges, and review labor.
Tool authority The model has no external side effects or is limited to read-only, sandboxed, low-risk tools. The model may call developer tools, search repositories, propose patches, or run bounded sandbox actions. The model supports planning or analysis for powerful workflows, but actions remain gated by policy and authorized operators. Least privilege, approval gates, audit logs, and rollback remain separate from model choice.
Data sensitivity Inputs are public, redacted, synthetic, or approved for the selected surface and processing mode. Inputs may include internal engineering or business material where workspace policy, role permissions, and regional requirements are satisfied. Inputs are sensitive enough that capability, verification, and governance matter more than unit price. Do not use model routing to bypass workspace policy, data-region rules, retention rules, or access controls.
Long context The long context is mostly stable reference material and the output is narrow enough to evaluate reliably. The model must reason across large files, repositories, specifications, or conversation history with meaningful tool use. The task requires the strongest available end-to-end reasoning across a large, mixed, or ambiguous context. Account for the long-context price multiplier above 272,000 input tokens for the full request.
Review burden Review is quick, structured, and cheaper than using a stronger model for every attempt. Review requires domain expertise, but Sol reduces rework or escalation compared with Luna in local tests. Review is expensive, scarce, or safety-critical, making a stronger first-pass answer valuable even at higher token cost. Include reviewer time and escalation cost in the decision, not only API tokens.
Reversibility Bad outputs can be discarded with no lasting effect. Bad outputs may create wasted work but can be reverted before production impact. Bad outputs could be difficult to unwind, trigger external obligations, or affect regulated, legal, financial, security, or health-related contexts. Require authorized human approval before publication, deployment, payment, account change, legal commitment, or destructive operation.

The table deliberately treats Luna as appropriate for many narrow tasks without treating Luna as the default for everything. A lower listed price can make Luna attractive for high-volume classification, extraction, summarization, or draft-generation workloads, but only after a representative evaluation confirms acceptable quality, escalation rate, latency, and review burden. If a cheap model produces more retries, longer outputs, more human corrections, or more failed tool calls, the total task cost can exceed a more capable route.

The same caution applies to Sol and Astra. A stronger model is not a license to skip approval, sandboxing, least privilege, logging, or rollback. OpenAI’s launch and model-reference materials identify relative positioning and capabilities, but they do not establish that any model can safely perform high-impact actions unattended in a particular enterprise, classroom, legal-technology workflow, security operation, or software delivery pipeline.

Compare the model facts that should be hard-coded into routing policy

Before a team writes a routing rule, it should make the official model facts explicit. OpenAI’s model references list the same 1,050,000-token context window, 922,000 maximum input tokens, and 128,000 maximum output tokens for GPT-6 Luna, GPT-6 Sol, and GPT-6 Astra. That shared context window does not mean the models have equal capability, equal latency, equal cost, or equal fit for difficult long-context reasoning.

Model Positioning in this guide Knowledge cutoff listed by OpenAI Context and output limits listed by OpenAI Supported reasoning efforts listed by OpenAI
GPT-6 Luna Focused high-volume work where local evaluation shows acceptable quality and review cost. May 18, 2026 1,050,000-token context window; up to 922,000 input tokens; up to 128,000 output tokens. none, low, medium default, high, xhigh, max.
GPT-6 Sol Complex coding, tool-using, agentic, and multi-step workflows where Luna is insufficient or review cost rises. April 20, 2026 1,050,000-token context window; up to 922,000 input tokens; up to 128,000 output tokens. none, low, medium default, high, xhigh, max.
GPT-6 Astra Hardest end-to-end work where the strongest overall GPT-6 model is justified by risk, complexity, or review economics. April 30, 2026 1,050,000-token context window; up to 922,000 input tokens; up to 128,000 output tokens. low, medium, high, xhigh, max. OpenAI’s cited Astra page does not list none.

Luna’s later published knowledge cutoff does not make it the right choice for all current-facts work. A cutoff date is not a freshness guarantee, a source-verification mechanism, or a quality ranking. Current legal, regulatory, financial, security, product, and medical-adjacent information still requires retrieval from approved dated sources, citation review, and human validation where appropriate. A model with an earlier cutoff plus reliable retrieval and review may outperform a later-cutoff model used without evidence.

Teams should also avoid copying the same reasoning configuration across all three models. Sol and Luna list none reasoning effort, while Astra does not list none in the cited model reference. OpenAI’s reasoning documentation states that supported effort values are model-dependent. Routing code should validate the selected model, effort, and mode before request execution rather than assuming that every GPT-6-family configuration is portable.

Use exact API prices, then convert them into task prices

OpenAI’s model pages publish Standard API token prices per one million tokens. These list prices are necessary for budgeting, but they are not the same as total workload cost. A production request can include uncached input, cache writes, cached reads, visible output, hidden reasoning tokens billed as output, retries, tool charges, processing-mode multipliers, regional premiums, and human review time.

Model Input Cached input read Cache write Output
GPT-6 Luna $0.10 per 1M tokens $0.01 per 1M tokens $0.125 per 1M tokens $0.50 per 1M tokens
GPT-6 Sol $2.00 per 1M tokens $0.20 per 1M tokens $2.50 per 1M tokens $10.00 per 1M tokens
GPT-6 Astra $10.00 per 1M tokens $1.00 per 1M tokens $12.50 per 1M tokens $50.00 per 1M tokens

The cache-write and cached-read prices follow OpenAI’s prompt-caching economics for GPT-5.6-and-later models: cache writes cost 1.25 times the uncached input rate and cached reads cost 0.1 times the uncached input rate. A cache write is therefore not automatically cheaper than ordinary input processing. It becomes economical only when a stable prefix is reused enough times and when the rest of the request does not erase the savings through extra output, reasoning, retries, tools, or review.

Requests above 272,000 input tokens need special treatment. OpenAI documents that, for requests above that threshold, input and cache rates are multiplied by 2x and output rates by 1.5x for the full request. This rule can materially change the ranking of candidate routes. A Luna request that crosses the long-context threshold, generates extensive reasoning and output, and then requires expert correction may be more expensive than a shorter, better-targeted Sol request or an Astra request reserved for only the unresolved cases.

Pricing variable Operational meaning Routing consequence
Uncached input New prompt content, instructions, files, images, or conversation turns that do not reuse an eligible prefix. Large unstable prompts reduce the value of caching and make per-request design more important than model price alone.
Cache write A new eligible prefix written at 1.25x the uncached input rate. Use for stable shared prefixes that are expected to be reused; do not prewarm sensitive or unapproved material merely to chase discounts.
Cached input read Eligible reused prefix billed at 0.1x the uncached input rate. Helpful for repeated long context, but it does not validate source truth or authorize cross-user data sharing.
Output and reasoning tokens Visible output and hidden reasoning tokens are billed as output tokens; reasoning tokens also count against output and context limits. Higher reasoning effort can improve hard-task performance but can increase cost and latency. Measure by task class.
Processing mode OpenAI’s cited model pages state Batch and Flex are priced at 50% of Standard rates, while Fast mode is 2x applicable rates. Compare cost only after matching the actual processing mode used by the workflow.
Regional processing The cited model pages state regional processing adds a 10% premium where available for Sol and Luna, and EU data residency for Sol and Luna is available only with Standard processing. Do not trade away regional, residency, or compliance requirements for a cheaper route.
Tool calls Tool calls may add separate charges and can also increase latency, review needs, and failure modes. Route tool-authorized tasks by capability and governance, not just token price.

For ChatGPT Work and Codex users, API prices should not be treated as a direct substitute for plan allowances or workspace usage behavior. OpenAI’s Help Center notes that Sol, Luna, and Astra can consume usage differently depending on task, input and output size, reasoning settings, and Fast Mode. ChatGPT plan usage and API-key billing are separate systems. Administrators should therefore manage Work/Codex policy and API cost accounting as related but distinct controls.

Distinguish endpoint support from product-surface availability

A model can be available in one surface but not appropriate or available in another. OpenAI’s launch and Help Center materials state that Sol and Luna are available in ChatGPT Work and Codex for specified paid and organizational plans, and that Luna is available to Free and Go users in the desktop application. The launch page also states that, at launch, Sol and Luna are not yet available in ordinary Chat. Availability can vary by plan, role, workspace setting, rollout, and policy.

In ChatGPT Work and Codex, a workspace starting default does not grant access to a model unavailable to a user’s role. Owners and administrators can configure starting model, reasoning level, speed, Fast Mode availability, and new-chat behavior for Work and Codex, but role-specific controls, browser and network controls, file permissions, and local or cloud permissions remain separate. A lower-cost model route must never expand a user’s access to files, repositories, tools, or external systems.

Capability or endpoint Luna Sol Astra Decision note
API model identifier gpt-6-luna gpt-6-sol Use the official model ID from the cited Astra model reference in implementation. Do not invent aliases or assume that ChatGPT picker names equal API identifiers.
Supported primary API surfaces listed by OpenAI Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. OpenAI recommends Responses for reasoning models and for built-in tools and general function-calling workflows.
Unsupported endpoints or features listed by OpenAI No Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions support on the cited page. No Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions support on the cited page. No Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions support on the cited page. Do not design a Sol, Luna, or Astra workflow around unsupported endpoints.
Chat Completions function calling limitation Supported only when reasoning_effort is none. Supported only when reasoning_effort is none. Validate against the Astra model reference and the selected API surface. Use Responses for built-in tools and normal function-calling workflows instead of depending on a narrow Chat Completions configuration.
Responses tools listed for the family Web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Tool availability does not equal permission to execute side effects. Bind every tool to least privilege and approval rules.

The endpoint differences matter most when teams attempt a simple “replace the model ID” migration. Switching from a prior model to Luna, Sol, or Astra while leaving old endpoint assumptions, tool wiring, reasoning settings, retry behavior, and cache policy unchanged can produce silent regressions. Migration rollback should restore the model selection, prompt/cache policy, tool availability, state-handling behavior, and prior validated controls together, not just flip a single model field.

Choose reasoning effort as a cost, latency, and reliability lever

OpenAI’s reasoning documentation states that GPT-6 Sol and Luna default to medium effort, and that supported effort values are model-dependent. GPT-5.6 and GPT-6 models support standard and pro reasoning modes; mode and effort are independent. Pro mode performs more model work, increases token usage and latency, and bills those tokens at the selected model’s standard rates. It should be evaluated as a deliberate tradeoff, not enabled by habit.

Reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible. If max_output_tokens is too low, a response can end as incomplete, potentially before any visible answer is produced. Applications must explicitly handle incomplete responses by failing closed, retrying under controlled conditions, escalating to review, or returning a clear non-result. They should not treat missing visible text as a successful answer.

Task pattern Reasoning-effort starting point Escalation trigger Cost and safety warning
Simple extraction, formatting, or classification from approved inputs Test Luna at none, low, or medium where supported and appropriate. Escalate if accuracy, consistency, or structured-output validity falls below the local threshold. Do not use low effort for high-impact decisions merely because labels are short.
Drafting internal summaries or first-pass knowledge-work outputs Test Luna or Sol at medium. Escalate if reviewers spend significant time correcting omissions, unsupported claims, or tone and policy failures. Require source verification for current facts and sensitive claims.
Repository analysis, test generation, refactoring proposals, or patch review Test Sol at medium or high, with Astra reserved for hard failures or critical reviews. Escalate if the model misreads dependencies, proposes unsafe patches, or fails tool-selection tests. Generated code needs human review, tests, and deployment controls before merge or release.
Ambiguous multi-step planning with tools Test Sol at higher efforts and compare Astra for the hardest cases. Escalate when tool authority, uncertainty, or downstream consequence rises. Powerful tools require explicit policy, least privilege, audit logging, and approval before side effects.
Regulated, legal, financial, health-adjacent, security, or external-commitment workflows Use the model only for bounded assistance; do not treat reasoning effort as a substitute for qualified review. Escalate to authorized human decision-makers for any consequential recommendation or action. No model choice provides personalized legal advice, regulated approval, or authority to act on behalf of the organization.

On supported GPT-6 models, OpenAI documents that an appended configuration_update can change reasoning effort during a conversation while preserving the earlier cacheable prefix. This can help teams avoid rewriting stable context solely to increase or decrease effort. However, changing top-level reasoning settings can affect cache reuse, and support varies by model. Production code should check the model documentation and record effort, mode, and cache state in evaluation logs.

{
  "model_route_policy": {
    "task_class": "repository_patch_review",
    "default_candidate": {
      "model": "gpt-6-sol",
      "reasoning_effort": "medium",
      "mode": "standard"
    },
    "escalate_when": [
      "security_sensitive_files_changed",
      "tests_fail_or_are_missing",
      "tool_arguments_are_uncertain",
      "reviewer_disagreement",
      "incomplete_response",
      "output_requires_external_commitment"
    ],
    "escalation_candidate": {
      "model": "gpt-6-astra",
      "reasoning_effort": "high",
      "mode": "standard"
    },
    "human_approval_required_for": [
      "merge",
      "deployment",
      "permission_change",
      "destructive_action",
      "external_message"
    ]
  }
}

The sample policy above is an implementation pattern, not an OpenAI product guarantee. It shows how to separate default routing from escalation and approval. A real organization should replace the task class, thresholds, tool names, and review rules with locally validated criteria, and should never include secrets, private keys, personal records, or unapproved confidential data in routing logs.

Why Luna’s price advantage is real but not decisive

Luna’s listed Standard API prices are dramatically lower than Sol and Astra: $0.10 per million input tokens and $0.50 per million output tokens, compared with Sol at $2 and $10, and Astra at $10 and $50. That price difference is operationally important for high-volume work. It can allow more shadow tests, broader coverage of low-risk tasks, and cheaper first-pass processing when the quality and review profile is acceptable.

The decision failure is assuming that the cheapest token is the cheapest task. A Luna answer that takes three attempts, emits a long uncertain response, fails a structured-output schema, calls the wrong tool, or forces a senior reviewer to rewrite the result may cost more than a Sol answer that succeeds once. In high-risk settings, the price of a wrong action, late escalation, or missed defect can dominate token spend entirely.

Reason Luna may appear best Why that can be misleading Decision rule
Lowest listed token prices Total task cost includes output, reasoning tokens, cache writes, cached reads, retries, tools, long-context multipliers, processing mode, regional premiums, and review labor. Compare completed task cost at the same acceptance threshold, not raw input price.
Later published knowledge cutoff than Sol or Astra A cutoff date is not a guarantee of currentness, accuracy, source quality, or domain suitability. Use retrieval, dated sources, and human verification for current or consequential facts.
Same listed context window as Sol and Astra A shared window does not imply equal long-context comprehension, reasoning, latency, or error behavior. Test long-context cases with representative documents and real review rubrics.
Availability in Work and Codex for many users Availability depends on plan, role, workspace setting, rollout, and policy; a default does not grant access. Confirm surface, role, and workspace policy before standardizing a route.
Good enough for many routine outputs Routine-looking tasks can hide sensitive data, external commitments, or irreversible tool actions. Classify by error cost and authority before considering volume discounts.

A conservative routing policy should let Luna win where it actually wins: high-volume, low-authority, easy-to-review tasks with stable prompts and clear acceptance tests. It should not let Luna inherit work merely because the request is short, the user is impatient, or the budget owner wants a lower headline rate. The cheapest model should still lose automatically when the action is consequential, the evidence is weak, the tool authority is broad, or the failure is hard to reverse.

Map Work and Codex permissions before model defaults

In ChatGPT Work and Codex, administrators can configure starting model, reasoning level, speed, Fast Mode availability, and new-chat behavior, according to OpenAI’s Help Center. Those settings are useful for steering behavior, but they are not a complete access-control system. Role permissions, browser and network controls, file permissions, local and cloud permissions, and workspace policies must be handled separately and verified during rollout.

A Work or Codex default should be treated as a starting preference, not a security boundary. If a user does not have access to a model, setting that model as a default does not grant access. If a user does have access to a model, that access does not automatically authorize broader repository access, external network use, local file reads, cloud execution, or tool-mediated side effects. Model selection and tool authority should be reviewed together.

Administrator decision Risk if treated as only a model setting Recommended operational check
Set Luna as starting model for high-volume teams Teams may use Luna for tasks with hidden legal, customer, or security consequences. Pair the default with task-class guidance, escalation rules, and restricted tool authority.
Set Sol as starting model for engineering or Codex-heavy teams Users may assume model strength replaces tests, code review, or deployment controls. Require repository permissions, sandboxing, review, and CI checks before merge or release.
Reserve Astra for expert workflows Astra may be overused as a prestige default or underused when risk justifies it. Define escalation criteria based on complexity, error cost, review burden, and reversibility.
Enable faster modes for responsiveness Fast mode can change effective cost and may encourage premature acceptance of outputs. Measure cost, latency, and quality under the exact mode and require review for consequential work.
Allow tool use in Codex workflows A model may be able to propose or execute actions beyond what the user intended. Use least privilege, explicit approvals, audit logs, and rollback for tool-authorized workflows.

For enterprise administrators, the right question is not “Which model should everyone use?” but “Which model, effort, tool authority, and review rule should apply to this task class for this role in this workspace?” A single default may be convenient, but mature governance usually requires a small number of documented routes: routine low-risk work, engineering analysis, sensitive review, and high-impact escalation.

Use long context only when it improves the decision

All three models list a 1,050,000-token context window, but large context is not free operationally. Long prompts can increase latency, raise cost, complicate review, and introduce irrelevant or conflicting material. Above 272,000 input tokens, OpenAI’s cited model pages apply the long-context multipliers to the full request: 2x input and cache rates, and 1.5x output rates. A routing policy should therefore ask whether long context is necessary, not merely whether it is supported.

Long context is most defensible when the model must compare multiple source documents, inspect a repository, reconcile a contract and policy set, or maintain continuity across a complex technical investigation. It is least defensible when the prompt includes every available file because no one built retrieval, chunk selection, or source triage. The best model route can be a smaller, cleaner context sent to Sol or Astra rather than a huge, noisy context sent to Luna.

Long-context pattern Model-routing implication Review implication
Stable policy, style, or coding standards reused across many requests Consider prompt caching and evaluate Luna or Sol depending on task risk. Review the stable prefix for correctness and authorization before reuse.
Large repository or multi-file coding task Start with Sol for bounded engineering workflows; reserve Astra for difficult architectural or high-risk reviews. Require tests, code-owner review, and deployment gates before merge.
Large legal, compliance, or policy corpus Do not choose by price alone; use retrieval, citations, and qualified human review. Outputs are assistance, not personalized legal advice or final regulatory determinations.
Current-facts research over many sources Knowledge cutoff should not decide the route; retrieval quality and source verification matter more. Require dated citations and flag unsupported, contradictory, or stale evidence.
Conversation history accumulated over time Compaction or rewriting earlier turns can affect cache reuse and may change model behavior. Summaries should be reviewed when they carry decisions, assumptions, or commitments forward.

Prompt caching can reduce repeated computation for stable prefixes, but it is not semantic memory, source validation, or permission management. A cached prefix can include developer messages, tool definitions, conversation history, text, images, documents, and supported audio where supported, but cache reuse depends on compatible settings and unchanged rendered prefixes. Teams should preserve stable prefixes only when doing so is operationally and semantically correct.

Set escalation and fallback rules that fail closed

A good decision matrix is incomplete without escalation and fallback behavior. The application should know when to move from Luna to Sol, from Sol to Astra, or from any model to human-only handling. Escalation should be based on task evidence: incomplete responses, uncertainty, schema failure, missing citations, tool-call mismatch, safety-sensitive content, reviewer disagreement, production impact, or irreversible action.

Recommended escalation sequence for a bounded workflow:

1. Try the lowest-risk approved route for the task class.
2. Validate output against structured requirements, source requirements, and tool-policy rules.
3. If validation fails, retry only when the retry is safe, bounded, and logged.
4. Escalate to a stronger model when local evidence shows capability is the limiting factor.
5. Escalate to an authorized human when consequence, uncertainty, or authority exceeds the automation boundary.
6. Block external side effects until approval, audit logging, and rollback conditions are satisfied.

This sequence is a recommendation, not a claim about built-in product behavior. It is intentionally conservative because the model that writes the answer should not be the only control deciding whether the answer is safe to act on. In enterprise, education, legal-technology, security, and software delivery settings, a valid output still needs policy review when it affects real people, systems, money, rights, accounts, grades, employment, legal obligations, or production infrastructure.

Escalation signal Likely next step Do not do this
incomplete response status or no visible output Fail closed, retry with a safe configuration, or escalate to review. Do not accept an empty or partial answer as success.
Missing or unverifiable current facts Use approved retrieval, require dated sources, and escalate if the evidence remains weak. Do not rely on knowledge cutoff as proof of freshness.
Tool call selected incorrectly or arguments are uncertain Block the tool call, request human review, or run in a sandbox without external side effects. Do not let a cheaper route expand tool authority.
Reviewer spends too long correcting the output Measure review time and test Sol or Astra against the same cases. Do not optimize only the model invoice while hiding labor cost.
Task becomes legally, financially, medically, security, or safety consequential Move to qualified human decision-making with model assistance only where approved. Do not treat model strength or high reasoning effort as professional authorization.
Destructive, external, or irreversible action is requested Require explicit authorized approval and rollback planning before action. Do not allow autonomous deletion, payment, booking, publication, permission change, or deployment.

The safest fallback is sometimes not a stronger model but a narrower task. Instead of asking Luna to summarize a sensitive contract and recommend a course of action, the workflow might ask it to extract non-sensitive clause headings for human triage. Instead of asking Sol to deploy a patch, the workflow might ask it to identify tests and propose a diff for review. Instead of asking Astra for a final legal conclusion, the workflow might ask it to organize issues and cite the source passages that a qualified professional must examine.

Translate the matrix into four route templates

Teams can make the decision matrix operational by defining route templates. Each template should specify the task class, approved surfaces, model candidates, reasoning configuration, tool authority, data class, review rule, and rollback behavior. The examples below are policy patterns that organizations can adapt after representative evaluation; they are not universal deployment instructions or claims about guaranteed model performance.

Route template Candidate model Good fit Required review and controls
Routine low-authority route Luna Approved summaries, labels, structured extraction, internal drafts, and repetitive transformations with easy verification. Sampling review, schema validation, source checks where facts matter, and escalation for sensitive or consequential content.
Engineering analysis route Sol Repository questions, code explanation, test generation, refactoring proposals, bounded Codex workflows, and tool-assisted debugging. Sandboxing, least-privilege repository access, tests, code review, and human approval before merge or deployment.
Hard-work escalation route Astra Ambiguous, high-value, multi-document, cross-functional, or high-review-cost work where strongest overall capability is justified. Expert review, audit trail, source verification, and explicit approval before external or irreversible actions.
Human-only or model-assisted advisory route No autonomous model route Legal commitments, regulated determinations, payments, account changes, sensitive employment decisions, safety-critical operations, destructive actions, and final publication approvals. Qualified human decision-maker; model may assist with drafting, issue spotting, or summarization only within approved boundaries.

A route template should record why a model was selected and what would cause reassignment. For example, a support-triage team might approve Luna for topic classification until sampled precision falls below a threshold, sensitive categories rise, or reviewer correction time exceeds the cost of Sol. An engineering platform team might approve Sol for patch suggestions but require Astra review for security-critical files or large architectural migrations. A legal-operations team might use Astra for issue organization but keep final advice and client communications under attorney review.

The most useful policy language is specific and testable. “Use the best model for important work” is not enforceable. “Use Luna only for approved low-impact extraction tasks; escalate to Sol after two schema failures or when the request includes repository tool use; escalate to Astra for security-sensitive architecture reviews; require authorized approval before merge, deployment, publication, external message, payment, or destructive action” gives administrators, developers, and reviewers a rule they can implement and audit.

Define the minimum evidence before changing a default

OpenAI’s evaluation best-practices guidance recommends task-specific objectives, representative data, metrics, comparisons, and continuous evaluation. For Sol, Luna, and Astra selection, the minimum evidence should include typical cases, edge cases, adversarial cases, tool-selection cases, tool-argument cases, and human-review calibration. Public benchmarks and launch examples can inform hypotheses, but they do not determine whether a model is fit for a specific Work, Codex, or API-backed workflow.

Evidence item What to measure Why it affects model choice
Task success rate Whether the output satisfies the local rubric, structured format, factual requirements, and policy constraints. A cheaper model is not cheaper if it fails acceptance more often.
Reviewer correction time Minutes spent verifying, editing, rejecting, or escalating each output. Human labor can dominate token cost for knowledge work, legal-tech, security, and coding workflows.
Tool correctness Tool selection, argument accuracy, side-effect prevention, and appropriate refusal to use tools. Tool-authorized tasks need stronger controls than text-only tasks.
Cost per accepted output Input, cached input, cache write, output, reasoning tokens, retries, tool charges, processing mode, regional premiums, and review labor. Token price alone can misrank models.
Latency distribution Median and tail latency under real context size, effort, mode, tools, and routing policy. Interactive Work and Codex use can fail operationally even when final quality is acceptable.
Escalation and fallback rate How often Luna escalates to Sol, Sol escalates to Astra, or any route escalates to human-only handling. A route that escalates too often may be the wrong default.
Incomplete response handling Frequency of incomplete status and whether the application safely detects it. Reasoning tokens can consume output budget before visible content appears.

Administrators should not change a workspace default or production route without a rollback bundle. The rollback should restore the previous model, reasoning configuration, prompt and cache behavior, tool definitions, permission assumptions, state handling, monitoring alerts, and review instructions. Switching only the model ID can leave incompatible assumptions in place, especially when endpoint support, function-calling behavior, or reasoning settings differ across routes.

Operational rule: choose the lowest-cost route that passes the local quality, safety, permission, review, and reversibility requirements for the task class. If the task is consequential, externally visible, regulated, destructive, permission-changing, payment-related, or difficult to reverse, authorized human review is part of the route rather than an optional afterthought.

This decision matrix should be revisited whenever the workflow changes: new tool access, new file classes, new regions, new workspace roles, new latency requirements, new legal or compliance obligations, new repository permissions, new retrieval sources, or new model documentation. Model choice is not a one-time procurement decision. It is an operating policy that must preserve evidence, controls, and human accountability as the work changes.

Control effort, permissions, and escalation as separate operating levers

Choose GPT-6 Sol, Luna, or Astra for ChatGPT Work and Codex: Capability, Cost, Reasoning Effort, Permissions, and Review — second editorial workflow visual

Model choice is only one control in a Work, Codex, or API-backed workflow. The same task can become cheap and safe, slow and over-permissioned, or dangerously under-reviewed depending on reasoning mode, effort, function-calling interface, tool policy, workspace role, local-file access, network access, and review gates. OpenAI’s documentation separates these dimensions: Sol and Luna default to medium reasoning effort, GPT-5.6 and GPT-6 models support standard and pro reasoning modes, Work and Codex defaults do not override role access, and ChatGPT plan usage is separate from API-key billing.

The practical rule is to route by measured task behavior and consequence, not by model prestige. A high-volume summarization or triage workflow that passes local evaluation on Luna should not be escalated to Sol or Astra merely because a larger model exists. A code-modification workflow that repeatedly fails tests, selects unsafe tools, or proposes irreversible changes should escalate in capability, effort, review, or permissions only under a documented rule. If the consequence is high, escalation may mean “stop and require a human,” not “give the model more authority.”

Use standard and pro mode deliberately, not as a quality slogan

OpenAI’s reasoning documentation states that GPT-5.6 and GPT-6 models support standard and pro reasoning modes, and that mode and effort are independent. Pro mode performs more model work, increases token usage and latency, and bills those tokens at the selected model’s standard rates. That means pro mode is not a free accuracy switch; it is a request for additional computation that must be justified by evaluation results, incident history, or task consequence.

For operational planning, treat standard mode as the baseline for routine evaluated work and pro mode as a controlled escalation path. A team might allow standard mode for drafting migration notes, generating unit-test candidates, or classifying support tickets, while reserving pro mode for tasks that combine ambiguous requirements, long dependency chains, architectural tradeoffs, or documented failure patterns under standard mode. The decision should be recorded with the task class, model ID, effort, tools, expected output, and required review.

Mode decision Appropriate use Operational warning Review expectation
Standard mode Routine tasks that pass local evaluations at an acceptable rate and have bounded consequences. Do not treat a successful sample as proof that all future production cases are safe. Use normal review for publication, code merge, customer messages, or external action.
Pro mode Harder cases where additional model work has been shown to improve task outcomes enough to justify added cost and latency. OpenAI states pro mode increases token usage and latency; it can also increase output-token billing because reasoning tokens are billed as output tokens. Keep human approval for consequential results; pro mode does not replace authorization.
Stop instead of escalate Cases involving missing authority, unclear instructions, regulated decisions, destructive actions, or conflicting evidence. A stronger model can still act on a flawed premise, wrong source, or excessive permission grant. Require an authorized human decision before continuing.

Reasoning tokens are not visible chain-of-thought. OpenAI states that reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible. Applications therefore need budget and completion handling that accounts for invisible reasoning work. If `max_output_tokens` is too low, a response can end with an `incomplete` status before any visible answer is produced, so production code must detect that status and avoid treating an empty or partial answer as a valid decision.

{
  "operational_rule": "Handle incomplete reasoning responses explicitly",
  "if_response_status": "incomplete",
  "do_not": [
    "publish the partial answer",
    "merge generated code",
    "send an external message",
    "retry with broader permissions automatically"
  ],
  "safe_next_steps": [
    "log model, effort, mode, tools, and token budget",
    "retry only within the same permission boundary if policy allows",
    "escalate to human review when the task is consequential",
    "record whether output limits, tool failure, or input size likely caused the stop"
  ]
}

Choose effort by task evidence, not by habit

OpenAI’s model references list different supported reasoning efforts for the three models. Sol and Luna support `none`, `low`, `medium`, `high`, `xhigh`, and `max`, with medium documented as the default. Astra supports `low`, `medium`, `high`, `xhigh`, and `max`; `none` is not listed for Astra. Routing code must validate the target model’s supported settings instead of copying one configuration across the family.

The presence of `none` for Sol and Luna is useful for simple transformation, extraction, or function-calling cases where extra reasoning is not needed, but it is not a guarantee of correctness. A low-effort or no-effort request can still produce an answer that needs source verification, schema validation, and permission checks. Conversely, `max` effort should not be used as a default for all hard-looking tasks because reasoning tokens are billed as output tokens, latency can increase, and over-budget responses can become incomplete.

Effort setting Supported models in cited references Good candidate tasks Decision rule
none Sol and Luna; not listed for Astra Constrained extraction, formatting, simple routing, or compatible Chat Completions function-calling cases. Use only after schema checks and evals show acceptable results; do not apply to Astra.
low Sol, Luna, and Astra Short, low-consequence work where the expected answer is direct and easy to verify. Escalate if failures cluster around multi-step reasoning, tool selection, or ambiguity.
medium Sol, Luna, and Astra; default for Sol and Luna General Work and Codex tasks after local evaluation confirms acceptable behavior. Use as the baseline until task-specific data justifies lowering or raising effort.
high or xhigh Sol, Luna, and Astra Complex coding, analysis, planning, and tool-use cases with measurable improvement from added reasoning. Require metrics showing fewer material failures or lower review burden, not just longer answers.
max Sol, Luna, and Astra Exceptional tasks with high complexity, strong review gates, and explicit cost/latency acceptance. Prefer bounded use with logging, output limits, and a human checkpoint before action.

For API conversations, OpenAI documents that an appended `configuration_update` can change reasoning effort on supported GPT-6 models while preserving the earlier cacheable prefix. That matters for long-running workflows that have a stable system prompt, tool definitions, and accumulated context. Changing top-level reasoning settings can affect cache-sensitive configuration, so teams should prefer the documented update pattern where applicable and still confirm support in the relevant model documentation.

Prefer Responses for normal tool use and validate Chat Completions exceptions

OpenAI recommends the Responses API for reasoning models and for built-in tools and normal function calling. The cited model references state that Sol, Luna, and Astra support Responses, Chat Completions, Batch, streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. They also state that these models do not support the Assistants, Realtime, Live, fine-tuning, embeddings, speech, transcription, or legacy Completions endpoints on the cited pages.

The important exception is narrow: Sol and Luna support Chat Completions function calling only when `reasoning_effort` is `none`. Astra does not list `none`, so a configuration copied from a Luna or Sol Chat Completions function-calling path can fail or behave incompatibly when routed to Astra. If your workflow depends on tools, structured output, or multi-step reasoning, the safer default is to design around Responses rather than building a routing layer that assumes identical Chat Completions behavior across the family.

{
  "routing_validation": {
    "model": "gpt-6-luna",
    "endpoint": "chat_completions",
    "function_calling": true,
    "required_reasoning_effort": "none",
    "warning": "Do not reuse this configuration for Astra; Astra does not list none effort in the cited model reference."
  },
  "preferred_general_tool_path": {
    "endpoint": "responses",
    "applies_to": ["gpt-6-luna", "gpt-6-sol", "gpt-6-astra"],
    "validation_required": [
      "model-specific effort support",
      "tool availability",
      "workspace or API authorization",
      "human approval gates for consequential actions"
    ]
  }
}

Endpoint support is not the same as permission to use a tool in your environment. OpenAI’s model references list supported Responses tools such as web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Your organization still needs a tool policy that defines which roles, workspaces, repositories, files, domains, and execution contexts are allowed for each task class.

Write tool policies as allowlists, not model-specific wishes

A tool policy should describe permitted actions, not merely the model selected. For example, “Luna may summarize approved internal documents” is incomplete because it does not say whether web access is allowed, whether files can be uploaded, whether the output can be sent externally, or whether the model can call a patch tool. A stronger policy says the task may read only approved workspace documents, may not browse the public web, may not modify files, and must route any customer-facing output to a human reviewer.

For Codex-style work, separate read authority, edit authority, execution authority, network authority, and merge authority. A model may be allowed to inspect a repository and propose a patch without being allowed to run networked commands, modify deployment configuration, or create a pull request. The difference matters because OpenAI’s launch and model pages describe capability, not your organization’s approval structure.

Tool dimension Low-risk policy example Higher-risk policy example Required control
File access Read a designated folder of approved documentation. Read broad repository or workspace files. Role-based access, data classification, and audit logs.
Code modification Generate a patch suggestion for review. Apply a patch in a working tree. Human code review before merge or deployment.
Shell or code execution Run local tests in a sandbox with no secrets. Execute commands that access network, credentials, or production resources. Sandboxing, command allowlists, secret isolation, and approval.
Web or external access Retrieve public documentation for citation verification. Submit forms, post messages, purchase services, or interact with accounts. Explicit human authorization before any external side effect.
Computer use Inspect a controlled test environment. Operate business systems, customer records, finance tools, or admin consoles. Least privilege, session monitoring, review gates, and rollback plan.

OpenAI’s prompt-caching guidance also affects tool policy. Stable tool definitions and schemas can improve cache reuse, and OpenAI recommends changing callability with `allowed_tools` or `tool_choice: none` rather than removing definitions when possible. That is a performance and cost pattern, not a permission shortcut. If a tool is disallowed for a task, the application must enforce that disallowance even if keeping a stable tool definition would improve cache behavior.

Understand Work and Codex defaults before changing them

OpenAI’s Help Center says Sol and Luna are models for ChatGPT Work and Codex and are not available in ordinary Chat conversations at the time of the cited update. It also states that availability depends on plan, workspace settings, role permissions, and rollout access. Owners and administrators can configure starting model, reasoning level, speed, Fast Mode availability, and new-chat behavior for Work and Codex, but a workspace starting default does not grant access to a model unavailable to a user’s role.

This creates a common administrative trap: setting Sol or Astra as a starting default is not the same as provisioning every user for that model, nor is it the same as authorizing broader local or cloud access. Codex preserves a manually selected model, and the desktop Work/Codex picker and defaults are separate from ordinary Chat defaults. OpenAI also notes that Codex CLI does not use the desktop slider, so administrators should not assume that a desktop speed or reasoning setting governs command-line usage.

A practical rollout inventory should list each surface separately: Work desktop, Codex desktop, Codex CLI, web or mobile surfaces where relevant, and API-backed internal tools. For each surface, record the permitted models, starting defaults, role eligibility, file access, browser or network controls, local/cloud permissions, logging, review gates, and rollback method. Without that inventory, teams can mistakenly validate one experience while leaving another experience governed by different defaults.

{
  "workspace_default_change_record": {
    "surface": "ChatGPT Work and Codex desktop",
    "starting_model": "gpt-6-sol or admin-selected product label where available",
    "reasoning_level": "workspace-configured value",
    "does_not_grant": [
      "model access unavailable to the user's role",
      "ordinary Chat access where the model is not available",
      "Codex CLI slider behavior",
      "local files, cloud files, browser, or network permissions"
    ],
    "required_checks": [
      "role eligibility",
      "workspace policy",
      "rollout access",
      "tool and file permissions",
      "human review gates"
    ]
  }
}

Separate local access, cloud access, and review authority

Local and cloud access should be evaluated as separate risk classes. A model that can read local files may encounter secrets, unpublished code, personal information, or regulated records if the user’s environment is not prepared. A model that can access cloud repositories or workspace documents may cross project, client, or classroom boundaries if folder-level or role-level controls are loose. Do not compensate for broad file access by selecting a supposedly safer model; reduce access first.

For enterprise administrators, the minimum safe pattern is least privilege plus review. Give the workflow only the repositories, folders, documents, tools, and network destinations needed for the task. Keep secrets out of model-visible contexts, use redacted or approved evaluation corpora, and require authorized human approval for deployment, external communication, account changes, permission changes, purchases, payments, publication, legal commitments, or destructive operations.

Educators and parents should apply the same separation at a smaller scale. A student-facing workflow can be allowed to explain a public reading, help outline an essay, or review syntax without being allowed to access private family files, impersonate a student in a submission system, or send messages to teachers. If a school or household uses desktop access, the adult administrator should define which folders and accounts are appropriate rather than relying on the model selection to enforce boundaries.

Legal-technology professionals should be especially conservative. A model may help classify public cases, draft issue lists, or compare clauses under supervision, but legal advice, filing decisions, privilege calls, client communications, settlement positions, and production responses require qualified human review. Model capability, larger context windows, and higher reasoning effort do not create legal authority or resolve professional-responsibility duties.

Design Luna-to-Sol-to-Astra escalation with measured failures

A robust escalation ladder begins with a task class and an evidence threshold. Luna is a sensible candidate for focused high-volume work where local evaluation shows acceptable accuracy, format compliance, and review burden. Sol is a candidate when Luna’s failures cluster around coding complexity, multi-step planning, tool selection, or agentic workflows and Sol demonstrably reduces material failures for the same task. Astra is reserved for the hardest end-to-end work where the additional capability is justified by task complexity, consequence, and measured improvement under the same review constraints.

Escalation should never mean removing safeguards. If a task escalates from Luna to Sol because Luna selected the wrong tool in 8 of 100 representative test cases, the Sol run should inherit the same or stricter tool policy, not broader network access. If a task escalates to Astra because Sol cannot resolve a cross-repository architectural plan, Astra should still be prevented from merging code, changing permissions, or contacting external systems without human approval.

Observed condition Escalation action What must not change automatically Evidence to record
Luna passes quality thresholds and review burden is acceptable. Keep Luna for the task class. Do not escalate for prestige or because another model has a higher list price. Pass rate, reviewer corrections, token cost, latency, incomplete responses, and tool errors.
Luna fails on multi-step reasoning or tool selection in representative cases. Test Sol at a higher or appropriate effort level under the same permissions. Do not grant broader tools or external action authority during the comparison. Pairwise outcomes, failure taxonomy, tool-call arguments, and human labels.
Sol improves quality but remains unreliable for high-consequence cases. Route to Astra only if local evaluation shows material benefit; otherwise stop for expert review. Do not assume Astra eliminates review, rollback, sandboxing, or audit requirements. Residual failure types, severity, reviewer time, cost, and decision rationale.
A model produces incomplete, contradictory, unverifiable, or overconfident output. Fail closed and require human review or additional evidence. Do not auto-publish, auto-merge, auto-send, or auto-pay. Status, missing evidence, citations or source gaps, and reviewer disposition.
Costs or latency exceed the task’s budget without quality improvement. De-escalate model, mode, effort, context size, or tool use after validation. Do not de-escalate review for consequential actions. Task-price breakdown, cached and uncached tokens, output tokens, retries, and long-context multipliers.

De-escalation is just as important as escalation. If Sol at high effort performs no better than Luna at medium effort for a structured extraction workload, route future cases to Luna and keep Sol for exceptions that match a documented failure pattern. If Astra resolves rare planning failures but costs too much for routine tickets, reserve it for manually approved escalations or cases that trigger specific confidence, complexity, or severity criteria.

Use review gates that match consequence, not model confidence

Review gates should be based on what the output can affect. A fluent answer about a public API migration may only need technical review before publication. A patch that changes authentication, billing, access control, data deletion, logging, cryptography, or deployment configuration requires a qualified reviewer and tests even if it was generated by Astra at high effort. A proposed customer message, legal filing, financial recommendation, health-related communication, or school disciplinary note requires the appropriate human authority before it leaves the organization.

OpenAI’s evaluation guidance recommends representative data, task-specific objectives, metrics, continuous evaluation, and human-calibrated scoring. Apply that guidance to review gates by converting failure classes into gates. If tool arguments are often wrong, require tool-call review. If citations are weak, require source verification. If code compiles but changes behavior unexpectedly, require regression tests and code-owner review. If the workflow touches regulated or sensitive domains, require a qualified domain reviewer rather than relying on model self-assessment.

  1. Define the action boundary. Decide whether the model is drafting, recommending, editing, executing, submitting, or changing permissions.
  2. Assign a consequence tier. Label the task as reversible internal work, externally visible work, customer-impacting work, regulated work, financial work, safety-related work, or destructive work.
  3. Set the review role. Require the appropriate code owner, administrator, teacher, clinician, attorney, compliance officer, manager, or other authorized reviewer where applicable.
  4. Require evidence. Keep the prompt version, model, mode, effort, tools, sources, tests, token usage, incomplete status, reviewer decision, and rollback path.
  5. Block side effects until approval. Do not send, merge, deploy, purchase, book, delete, publish, file, or change access without explicit authorization.

For knowledge workers, a practical review threshold is “Can this output embarrass, bind, charge, expose, delete, or mislead someone if wrong?” If yes, the output needs human review before external use. For developers, the equivalent threshold is “Can this change alter data, permissions, security posture, build systems, or production behavior?” If yes, require tests, diff review, and rollback readiness before merge.

Make escalation observable with a failure taxonomy

Teams cannot improve routing if every bad answer is logged as “model failed.” Use a compact failure taxonomy that distinguishes task misunderstanding, missing context, stale knowledge, retrieval failure, citation failure, schema violation, tool-selection error, tool-argument error, unsafe action proposal, hallucinated completion, incomplete response, excessive latency, excessive cost, and reviewer rejection. This taxonomy supports measured escalation from Luna to Sol or Astra and measured de-escalation when a smaller route is sufficient.

Failure type Likely first fix When to test stronger model or higher effort
Missing or stale facts Improve retrieval, require dated sources, or add source verification. Only after the same evidence is available to each model.
Schema or format violation Use structured outputs, validation, and retry within the same authority. When violations persist after clear schema and validation feedback.
Tool-selection error Tighten allowed tools, tool descriptions, and task routing. When representative tests show a stronger model selects tools correctly more often.
Unsafe action proposal Reduce authority, add policy checks, and require human review. Escalate to a model only for analysis; do not expand execution permission automatically.
Incomplete response Adjust output budget, split the task, or handle status safely. When incompleteness reflects reasoning complexity rather than a budget or context design problem.
Reviewer rejection for reasoning quality Improve prompt, examples, context, and effort setting. When the rejection pattern remains after prompt and data issues are addressed.

Public benchmarks can suggest where to look, but they cannot replace this failure evidence. OpenAI’s launch materials describe Sol and Luna improvements and Astra’s overall strength, while also noting that benchmark results are evaluation-specific and may differ from production ChatGPT because of system prompts, tools, and deployment differences. Treat public scores as hypotheses to test, not as permission to bypass local gates.

Keep cost controls attached to the same route definition

A route definition should include the model, mode, effort, endpoint, tool policy, context policy, processing mode, review gate, and cost envelope. The model references list token prices, cached-input prices, cache-write prices, long-context multipliers above 272,000 input tokens, regional processing premiums where available for Sol and Luna, Batch and Flex pricing, Fast mode pricing, and possible separate tool-call charges. ChatGPT plan allowances and API billing remain separate systems, so a Work or Codex setting should not be translated directly into API cost without separate measurement.

If a route uses very long context, the shared 1,050,000-token context window across Sol, Luna, and Astra does not make the route equally economical or equally effective on all models. OpenAI documents that requests above 272,000 input tokens trigger 2x input and cache rates and 1.5x output rates for the full request. This can make “just send everything” a poor default even when the context window allows it.

Cost controls should not override safety controls. A lower-cost route never authorizes broader action scope, cross-tenant data exposure, reduced review, or weaker audit logging. If budget pressure requires de-escalation from Astra to Sol or Luna, the workflow must be re-evaluated for quality and risk; it should not quietly continue as if only the invoice changed.

Operational workflow: route, review, escalate, and roll back

The following workflow is a recommended operating pattern, not an OpenAI product guarantee. It is designed to help teams use Luna, Sol, and Astra with explicit evidence, role boundaries, and review gates across Work, Codex, and API-backed applications.

  1. Start with task classification. Identify the task’s purpose, data sensitivity, tool authority, reversibility, user role, required sources, and external side effects.
  2. Select the smallest plausible route. Use Luna for focused high-volume work when evaluation supports it, Sol for complex coding or agentic work when Luna fails materially, and Astra for the hardest end-to-end work when local evidence justifies the cost and review burden.
  3. Choose mode and effort separately. Begin with the documented default or evaluated baseline, then raise or lower effort only when measured quality, latency, cost, or incomplete-response data supports the change.
  4. Validate endpoint and function-calling constraints. Prefer Responses for normal tool use; use Chat Completions function calling for Sol or Luna only when the documented `reasoning_effort: none` condition fits the task.
  5. Apply tool and access policy. Enforce allowed files, repositories, domains, tools, network access, local/cloud access, and execution boundaries before the model runs.
  6. Run validation and review. Check schema, citations, tests, tool arguments, permissions, and incomplete status; route consequential outputs to authorized humans.
  7. Log the evidence. Record model, mode, effort, endpoint, tools, prompt version, input/output token classes, cache state if available, latency, failure type, reviewer decision, and cost.
  8. Escalate or de-escalate by rule. Move from Luna to Sol to Astra, or back down, only when the failure taxonomy and evaluation results justify the change.
  9. Roll back as a bundle. Restore model selection, effort, prompt/cache policy, tool availability, state handling, review gates, and previous validated behavior together.

Rollback deserves special attention because changing only the model ID can leave incompatible assumptions behind. A route that used Luna with `reasoning_effort: none` for Chat Completions function calling cannot be safely replaced with Astra by editing a string, because Astra does not list `none` and the endpoint/tool behavior may need redesign. A route that depended on stable tool definitions for caching may also need a prompt/cache policy rollback if the new configuration changes tool order, schema, verbosity, or reasoning settings.

Decision rules that keep the model ladder honest

The most reliable governance rule is simple: escalate for measured failure or consequence, de-escalate for measured sufficiency, and stop for missing authority. Measured failure includes repeated schema violations, wrong tool calls, failing tests, unverifiable claims, reviewer rejection, incomplete outputs, or unacceptable latency/cost against a defined task objective. Consequence includes legal, financial, safety, privacy, security, educational, employment, customer, administrative, and production-system impact.

A recommended Luna-to-Sol trigger is a representative evaluation showing that Luna’s failures are not caused by missing context, poor retrieval, unclear prompts, or excessive permissions, and that Sol materially reduces the failure class under the same tool policy. A recommended Sol-to-Astra trigger is a smaller set of hard cases where Sol still fails on end-to-end reasoning, cross-system planning, or complex code changes, and where Astra’s improvement is worth the added cost, latency, and review burden. A recommended de-escalation trigger is sustained evidence that a lower-cost route meets quality and safety thresholds with the same or lower reviewer burden.

Operational warning: A stronger model may reduce some failure rates, but it does not authorize unattended consequential action. Keep human approval for external messages, submissions, payments, purchases, bookings, destructive actions, permission changes, publication, legal commitments, regulated decisions, campaign launches, deployment, and other high-impact operations.

The final decision should be documented in a route card that an administrator, developer, reviewer, and auditor can understand. The card should say why the model was selected, which effort and mode are allowed, what tools are permitted, what data can be accessed, what review is required, what metrics will trigger escalation or de-escalation, and how to roll back. If the route card cannot explain those points, the workflow is not ready for broader rollout.

Turn the decision framework into an operating model

A model policy is only useful if it tells people who owns each decision, when a request can proceed, what evidence must be logged, and how to stop the workflow when model output is incomplete or unauthorized. For ChatGPT Work, Codex, and API-backed workflows, the operating model should treat Luna, Sol, and Astra as governed routes rather than interchangeable names in a picker. OpenAI positions Luna for focused high-volume work, Sol for complex coding and agentic workflows, and Astra as the strongest overall model for the hardest end-to-end work, but those positions do not override workspace permissions, tool controls, data-region requirements, or human review duties.

The most important operational distinction is between model selection and action authority. A request may be safe to draft with Luna, safer to analyze with Sol, and worth escalating to Astra for final reasoning, yet still be prohibited from sending an email, merging code, changing permissions, publishing a document, buying a service, or making a regulated decision without an authorized human. This separation prevents the common failure mode where a stronger model is mistakenly treated as a broader permission grant.

Recommended RACI for Work, Codex, and API model governance

The following RACI is a recommended operating structure for organizations that need repeatable model choices across knowledge work, engineering, support operations, legal-technology workflows, education administration, and enterprise automation. It is not legal advice or a substitute for sector-specific compliance review; it is a practical allocation of responsibility that keeps consequential decisions under human control.

Governed activity Responsible Accountable Consulted Informed
Define eligible task classes for Luna, Sol, and Astra AI product owner or workflow owner Business unit leader Security, legal, compliance, data governance, engineering Workspace administrators and affected users
Configure ChatGPT Work and Codex starting defaults where available Workspace owner or administrator Workspace owner IT, security, training lead, department owners End users and help desk
Maintain API routing rules and model-specific settings Platform engineering or application team Engineering owner Security, FinOps, data governance, QA Support and incident response
Approve tools, file access, browser/network access, hosted execution, or local/cloud permissions Security engineering and system owner CISO or delegated security owner Legal, privacy, enterprise architecture, data owners Workflow operators
Design and maintain task-specific evaluations Evaluation lead, QA lead, or applied AI engineer Product or engineering owner Domain experts, security, compliance, operations Administrators and reviewers
Approve production model-route changes AI change manager or release manager Service owner Security, FinOps, legal/compliance, affected business owner Users, support, audit function
Review consequential outputs before external action Authorized human reviewer Business decision owner Legal, compliance, clinical, financial, educational, or HR specialist as applicable Requester and audit record owner
Investigate incidents, unsafe outputs, permission drift, or cost anomalies Incident response lead Security or service owner Engineering, legal, privacy, FinOps, vendor-management owner Leadership, affected users, audit stakeholders

The RACI should be recorded outside the prompt itself in a policy repository, runbook, or change-management system. Prompts can remind users that approvals are required, but prompts cannot prove that the approver has authority, that the requester has permission to disclose the data, or that the downstream action is lawful. Enterprise administrators should therefore pair model defaults with role-based access, audit logging, training, and documented escalation paths.

Approval matrix for model route, tool authority, and consequence

A practical approval matrix should combine three dimensions: the model route, the authority granted to the workflow, and the consequence of an incorrect output. OpenAI’s model documentation identifies model capabilities, supported modalities, reasoning effort options, token pricing, and endpoint support; it does not authorize a local workflow to act without review. The organization must decide when an output is draft-only, when it can be used internally after review, and when it is too consequential for automated execution.

Work type Typical route Allowed output state Required human control Operational warning
Summarizing approved internal notes for personal productivity Luna first; Sol if ambiguity or synthesis failure is observed Draft summary or checklist User verifies accuracy before sharing or relying on it Do not include unnecessary personal, regulated, or confidential material.
High-volume classification, tagging, or triage with low consequence Luna with calibrated thresholds and sampling review Suggested label or queue placement Human review for uncertain, novel, sensitive, or high-impact cases Public benchmarks do not establish local precision, recall, or bias behavior.
Complex code review, refactoring plan, or multi-file reasoning Sol first; Astra for high-risk architecture or unresolved failures Recommendation, patch proposal, or test plan Engineer reviews before merge, deployment, permission change, or destructive command Codex model choice does not replace repository permissions, CI, security review, or rollback.
End-to-end agentic workflow with tool use Sol or Astra depending on evaluation evidence and consequence Plan, staged tool proposal, or sandboxed execution result Approval before external messages, purchases, bookings, submissions, account changes, or production writes Least privilege and allowlisted tools are mandatory; lower price never expands action scope.
Legal, financial, HR, health, education, youth-safety, or compliance-sensitive analysis Sol or Astra only after domain-specific evaluation, or no model route if prohibited Research aid, issue list, draft, or evidence table Qualified professional or authorized decision owner approves any use Do not treat model output as personalized legal, medical, financial, or employment advice.
Public communications, policy statements, marketing campaigns, or executive material Luna for first drafts; Sol or Astra for complex synthesis Draft copy, risk review, or source checklist Authorized publication owner approves before release Fact freshness requires retrieval, dated sources, and verification; knowledge cutoff is not a guarantee.

Approval rules should be enforced at the workflow layer rather than left to user memory. For example, an API-backed workflow can mark outputs with statuses such as draft, needs_reviewer, approved_for_internal_use, and blocked. A Codex workflow can require pull-request review and CI success before merge. A Work workflow can require manual confirmation before content leaves the workspace. These controls remain necessary even when routing escalates from Luna to Sol or Astra.

Keep routing decisions auditable

A routing decision record is the minimum evidence needed to explain why a workflow selected Luna, Sol, or Astra, what settings were used, what permissions were available, and whether review occurred. The record should be created for production routes, canary cohorts, high-cost requests, high-impact outputs, and incidents. It should not store secrets, personal records, confidential source text, or unnecessary raw prompts when a redacted reference or hash is sufficient for audit and debugging.

Routing decision record template

The template below is an example implementation artifact. Adapt it to your logging, privacy, retention, and legal requirements. Avoid storing credentials, tokens, private keys, personal identifiers, protected records, or privileged material unless there is a documented, approved, and legally reviewed reason to do so.

{
  "route_record_version": "2026-09-governed-model-route",
  "request_class": "code_review_plan",
  "surface": "Codex or API-backed workflow",
  "requested_model": "gpt-6-sol",
  "fallback_models_allowed": ["gpt-6-astra"],
  "models_not_allowed": ["gpt-6-luna"],
  "reasoning": {
    "mode": "standard",
    "effort": "medium",
    "model_specific_validation": "passed"
  },
  "task_risk": {
    "error_cost": "medium",
    "external_side_effects": "blocked",
    "regulated_decision": false,
    "destructive_action": false,
    "publication": false
  },
  "tool_policy": {
    "allowed_tools": ["file_search", "code_interpreter"],
    "blocked_tools": ["computer_use", "hosted_shell", "external_submission"],
    "human_approval_required_before_tool_expansion": true
  },
  "data_policy": {
    "approved_data_category": "internal engineering material",
    "regional_or_residency_requirement": "apply organization policy",
    "secrets_allowed": false,
    "raw_personal_data_allowed": false
  },
  "evaluation_basis": {
    "eval_suite_id": "local-code-review-holdout-2026-09",
    "minimum_pass_threshold_met": true,
    "human_calibration_completed": true,
    "last_review_date": "YYYY-MM-DD"
  },
  "output_status": "draft_needs_human_review",
  "human_reviewer_role": "authorized_engineer",
  "decision": "allow_draft_generation_only",
  "rollback_bundle": "model_prompt_tools_state_v3",
  "notes": "No merge, deploy, permission change, or destructive command without human approval."
}

The field model_specific_validation matters because the GPT-6 model references list different supported reasoning efforts. Sol and Luna include none, low, medium, high, xhigh, and max; Astra does not list none. A fallback from Sol to Astra therefore cannot blindly reuse every setting. The router must validate the target model’s supported configuration before sending the request.

Routing records should also separate ChatGPT plan usage from API billing. OpenAI’s Help Center notes that ChatGPT Work and Codex availability depends on plan, role permissions, workspace settings, and rollout access, and that plan usage and API-key billing are separate. A user’s access to a Work or Codex model does not prove that an API route is available, priced the same way, or governed by the same administrative settings.

Fallback behavior that fails closed

Fallback should protect correctness, cost control, and authority boundaries. A safe fallback is not “try a different model until something answers.” It is a pre-approved sequence that preserves the task class, data boundary, tool policy, output status, and human-review requirement. If those cannot be preserved, the workflow should stop and ask an authorized person to choose the next step.

Failure condition Permitted fallback Required stop condition What must not happen
Luna output fails local quality threshold Escalate to Sol with the same or stricter tool permissions Stop if the task is high consequence or unresolved after escalation Do not publish, decide, or execute based on the failed Luna output.
Sol cannot complete a complex reasoning or coding task Escalate to Astra if authorized and evaluated for that task class Stop if model-specific settings, data policy, or cost limits cannot be satisfied Do not copy Sol-only settings to Astra without validation.
Astra route unavailable or blocked by policy Return to human review or defer the task Stop immediately for regulated, destructive, external, or contractual actions Do not downgrade to a cheaper model to bypass approval or access constraints.
Tool, schema, file, or permission mismatch Run without the tool only if the route is explicitly approved for draft-only output Stop if the missing tool is required for verification or safe execution Do not invent evidence, simulate tool results, or claim a check was performed.
Cost ceiling, long-context multiplier, or quota concern Ask the user to narrow the task, summarize approved material, or request approval for the higher-cost route Stop if narrowing would remove evidence needed for a safe answer Do not silently omit material that changes the decision.

The same principle applies in ChatGPT Work and Codex user interfaces. A workspace default can help users start in a sensible place, but OpenAI’s Help Center states that a starting default does not grant access to a model unavailable to the user’s role. Administrators should communicate that defaults are starting points, not guarantees of access or permission to use every feature.

Handle incomplete responses explicitly

OpenAI’s reasoning guide notes that reasoning tokens are billed as output tokens and count against output and context limits even though they are not visible. It also warns that max_output_tokens can end a response as incomplete, potentially before any visible answer is produced. Applications must therefore check response status and not assume that every model call returns usable text.

// Example workflow policy, not an official SDK contract.
function handleModelResult(result, route) {
  if (result.status === "incomplete") {
    return {
      output_status: "blocked_incomplete_response",
      user_message: "The model response did not complete. No external action was taken.",
      next_steps: [
        "Preserve the audit record without storing sensitive raw content unnecessarily.",
        "Check whether max output, context length, reasoning effort, or task scope caused truncation.",
        "Retry only if the route permits retry and the action remains draft-only.",
        "Escalate to a human reviewer before any consequential decision or external action."
      ],
      allowed_external_side_effects: false
    };
  }

  if (!result.visible_text || result.visible_text.trim().length === 0) {
    return {
      output_status: "blocked_no_visible_answer",
      user_message: "The model produced no usable visible answer.",
      allowed_external_side_effects: false
    };
  }

  return {
    output_status: route.requires_human_review
      ? "draft_needs_human_review"
      : "draft_for_user_verification",
    allowed_external_side_effects: false
  };
}

An incomplete response should not automatically trigger a higher-effort retry if the task is consequential, the original prompt contained sensitive data, or the retry would exceed a cost or data boundary. The safer procedure is to record the incomplete status, preserve enough non-sensitive metadata for diagnosis, and require a reviewer to decide whether to narrow the task, adjust output limits, change effort, retrieve missing evidence, or stop.

Monitor routes after launch, not just before approval

OpenAI’s evaluation guidance recommends continuous evaluation with representative data, edge cases, adversarial cases, tool selection, tool arguments, and agent handoffs. That guidance is especially important for model routing because task distributions change after launch. Users learn how to ask for more, administrators change defaults, codebases evolve, and integrations gain or lose permissions. A route that was safe for draft-only summarization can become unsafe if it is later connected to a tool that sends messages or modifies records.

Monitoring signals for administrators, security teams, and product owners

Monitoring should combine quality, safety, cost, access, and operational signals. No single metric proves that Luna, Sol, or Astra is the right choice. A stable route should show acceptable task outcomes, reviewer agreement, manageable escalation rates, predictable cost, controlled tool use, and clear evidence that human approval gates are working.

Signal category Examples to track Decision rule
Quality and correctness Human pass rate, pairwise preference, factual correction rate, coding test outcomes, citation verification failures Escalate or pause the route if failures cluster by task type, source type, repository, or user cohort.
Review burden Reviewer time, number of edits, repeated escalation reasons, unresolved uncertainty, approval backlog A cheap model route is not cost-effective if it shifts excessive work to reviewers.
Incomplete or truncated output incomplete status, no visible answer, output limit hits, retry frequency Investigate task scope, reasoning effort, output limits, and long-context usage before accepting the output.
Tool and permission behavior Blocked tool attempts, unexpected tool selection, argument errors, denied file access, local/cloud permission conflicts Keep tools allowlisted and suspend routes that repeatedly request unauthorized actions.
Cost and latency Uncached input, cached input where applicable, output and reasoning tokens, retries, processing mode, regional premiums, long-context multipliers Reprice by task, not by token list price alone, and include reviewer cost where material.
Access and policy drift Role changes, workspace default changes, model availability changes, API route changes, data-region requirements Revalidate the route whenever access, defaults, tools, or data boundaries change.
Safety and compliance Prompt-injection exposure, sensitive-data incidents, regulated-topic misuse, publication without approval, destructive-action attempts Fail closed, preserve evidence, and involve security, legal, privacy, or compliance owners as appropriate.

Monitoring should preserve evidence without expanding privacy risk. For many workflows, it is enough to log route identifier, model ID, settings, token counts, tool names, decision status, reviewer outcome, and redacted failure categories. Storing full prompts and outputs may be unnecessary or inappropriate when they contain confidential documents, personal data, regulated records, or privileged analysis.

Periodic re-evaluation schedule

A model route should have an expiration date. The re-evaluation interval depends on consequence, rate of change, and exposure to external data. High-volume low-risk tasks may need statistical sampling and monthly review. Coding-agent routes that touch production repositories may need review at each major repository, dependency, tool, or permission change. Legal, HR, financial, health, education, youth-safety, or compliance-sensitive workflows should be reviewed with qualified domain owners before any expansion of scope.

Trigger Required re-evaluation Owner
New model availability, model documentation change, or route default change Run representative local evals and validate model-specific settings before adoption AI product owner and engineering owner
New tool, file permission, browser/network access, or execution capability Security review, least-privilege check, tool-argument evals, and human approval gate test Security owner and system owner
Shift from draft-only to internal decision support Reviewer calibration, error-cost analysis, escalation thresholds, and audit trail validation Business owner and compliance owner
Shift from internal support to external message, submission, payment, purchase, booking, publication, or deployment Formal approval workflow with authorized human signoff and rollback plan Business decision owner
Cost anomaly, latency spike, retry increase, or incomplete-response cluster Token accounting, long-context review, output-limit review, route comparison, and stop-condition test FinOps and platform engineering
Security incident, privacy concern, prompt-injection finding, or unauthorized tool attempt Immediate route suspension or restriction, evidence preservation, incident review, and policy update Security incident lead

OpenAI’s evaluation documentation also notes that the current Evals platform is being deprecated, becoming read-only for existing users on October 31, 2026 and scheduled to shut down on November 30, 2026. Teams should avoid creating a new critical dependency on that retiring platform. A safer pattern is to maintain an application-owned evaluation harness, store representative test sets under approved data controls, calibrate automated scoring against human labels, and keep route decisions portable across tooling changes.

Manage exceptions without weakening the baseline

Exceptions are inevitable: an executive needs a faster draft, an engineer wants Astra for a difficult debugging session, a support lead wants Luna for a new triage queue, or a legal-technology team wants Sol to summarize approved public materials. The governance mistake is to let exceptions become undocumented defaults. Every exception should state the user, task class, model route, data category, tools, review requirement, time limit, and rollback condition.

Exception request template

The following example keeps exception review concrete. It asks for the minimum information needed to decide whether the exception is appropriate without requesting secrets, personal records, or unnecessary confidential content.

Exception request: GPT-6 route or permission change

1. Requested change:
   - Model route: Luna / Sol / Astra
   - Surface: ChatGPT Work / Codex / API-backed workflow
   - Reasoning mode and effort requested, if applicable:
   - Tools or permissions requested:

2. Business purpose:
   - Task class:
   - Expected duration:
   - Expected users or cohort:
   - Why the standard route is insufficient:

3. Data boundary:
   - Approved data categories:
   - Prohibited data categories:
   - Regional or residency requirement, if any:
   - Confirmation that no passwords, tokens, private keys, or unapproved personal/regulated records will be submitted:

4. Consequence boundary:
   - Draft-only / internal decision support / external action:
   - Human approver role:
   - Actions explicitly blocked:

5. Evidence:
   - Relevant local eval results:
   - Known failure modes:
   - Monitoring plan:
   - Rollback plan:

6. Expiration:
   - End date or event:
   - Owner responsible for review:

Exception approvals should be time-limited. If the same exception is requested repeatedly, treat it as a signal to revise the baseline route through normal governance rather than allowing permanent informal access. If an exception involves regulated decisions, youth-related processing, employment decisions, financial consequences, health-related use, legal commitments, publication, payments, purchases, bookings, destructive actions, or permission changes, an authorized human must remain in control and the organization should obtain qualified review before expanding the workflow.

Example task portfolios for Luna, Sol, and Astra

The portfolios below are examples, not universal rules. They translate the earlier decision framework into practical operating patterns for teams that use ChatGPT Work, Codex, and API-backed workflows. Each portfolio assumes that availability, plan access, role permissions, workspace settings, endpoint support, and model-specific settings have already been verified against current OpenAI documentation and local policy.

Portfolio A: High-volume knowledge work with low external consequence

A knowledge-work portfolio can start with Luna when the task is repetitive, evidence is supplied by the user, and the output remains draft-only. Examples include rewriting an approved internal note for clarity, producing meeting-action summaries from organization-approved material, generating first-pass tags for a low-risk internal knowledge base, or preparing a checklist for a human operator. The route should escalate to Sol when the work requires synthesis across conflicting materials, careful comparison, or multi-step planning.

  • Recommended starting route: Luna for first-pass drafts and classifications, with sampling review.
  • Escalation trigger: Ambiguity, conflict, missing evidence, high reviewer edit rate, or repeated uncertainty.
  • Human control: User verifies facts before sharing; authorized owner approves publication or external communication.
  • Do not automate: External messages, public posts, contractual statements, or regulated determinations.

Portfolio B: Software engineering and Codex-assisted development

A coding portfolio often belongs in Sol when the work involves multi-file reasoning, debugging, refactoring, test generation, or agentic steps inside Codex. Astra can be reserved for the hardest architecture, security-sensitive design review, or unresolved problems where local evaluations show material benefit. Luna can still be useful for low-risk explanation, documentation drafts, or simple issue summarization, but it should not be routed to complex repository-changing tasks solely because it is less expensive.

  • Recommended starting route: Sol for complex coding and agentic workflows; Luna for simple explanation or issue triage; Astra for the hardest reviewed cases.
  • Required controls: Repository permissions, tool allowlists, CI, code review, security checks, and rollback.
  • Human control: Engineer approval before merge, deployment, destructive command, dependency change, credential rotation, or permission modification.
  • Operational warning: A model-generated claim that tests passed is not proof; verify with actual test evidence.

Portfolio C: Enterprise analysis, policy, and executive work

Enterprise analysis can use Luna for outline generation and basic summarization, Sol for structured analysis across multiple approved sources, and Astra for difficult end-to-end reasoning where the output informs high-stakes strategy. The operating risk is not only hallucination; it is also misplaced authority. A polished answer can still be wrong, outdated, overconfident, or based on materials that the requester was not allowed to combine.

  • Recommended starting route: Luna for drafts, Sol for analytical synthesis, Astra for the hardest reviewed strategy work.
  • Required evidence: Source list, date of sources, assumptions, unresolved questions, and reviewer notes.
  • Human control: Executive, legal, finance, HR, or compliance owner approves consequential use.
  • Operational warning: Knowledge cutoff dates are not freshness guarantees; retrieve and verify current facts.

Portfolio D: Legal-technology, compliance, and regulated support

Legal-technology and compliance portfolios should treat GPT-6 models as drafting, search, comparison, and issue-spotting aids unless a qualified authority approves a narrower use. Model output should not be represented as personalized legal advice, a final compliance determination, or a substitute for professional judgment. Sol or Astra may be appropriate for difficult analysis after local evaluation, but the final decision must remain with an authorized human.

  • Recommended starting route: Sol or Astra only for approved task classes; Luna only for low-risk formatting or summarization when permitted.
  • Required controls: Privilege review, data-minimization rules, source verification, retention policy, and reviewer signoff.
  • Human control: Qualified professional or accountable decision owner approves final use.
  • Operational warning: Do not submit privileged, regulated, or personal material unless the organization has approved that use and configured appropriate controls.

Portfolio E: Education, youth-facing, and parent-supported workflows

Education and youth-related workflows require extra caution because the consequence may involve minors, learning assessment, discipline, accessibility, or wellbeing. Luna may support simple practice generation or teacher drafting with approved materials. Sol can help design rubrics or compare curricular sources. Astra may be used for complex planning only where the institution has evaluated the workflow and retained educator control. No model should independently decide discipline, accommodations, grades, eligibility, or safety interventions.

  • Recommended starting route: Luna for low-risk drafts, Sol for educator-reviewed analysis, Astra only for difficult institution-approved planning.
  • Required controls: Age-appropriate use policies, data minimization, parental or institutional rules where applicable, and educator review.
  • Human control: Educators, administrators, parents, or qualified support professionals retain decision authority.
  • Operational warning: For crisis, self-harm, abuse, or urgent safety concerns, seek qualified real-world support through the appropriate institutional or emergency process.

Final decision rules for a durable model policy

The durable rule is simple: choose the cheapest model that has passed local evaluation for the task, but only inside the permissions, data boundaries, tool limits, review gates, and cost envelope approved for that workflow. Luna can be the right answer for focused high-volume work. Sol can be the right answer for complex coding and agentic workflows. Astra can be the right answer for the hardest end-to-end tasks. None of those routes authorizes unattended consequential action.

Before changing a default in ChatGPT Work or Codex, administrators should verify that the model is available for the relevant plan, role, workspace, and rollout state. Before changing an API route, engineers should verify endpoint support, model-specific reasoning efforts, token pricing, long-context behavior, tool support, output limits, and fallback settings. Before expanding tools or permissions, security teams should confirm least privilege, auditability, data boundaries, and incident response. Before relying on the output, reviewers should confirm evidence, uncertainty, and task-specific evaluation results.

A mature GPT-6 operating model does not ask “Which model is best?” in the abstract. It asks which model, at which effort and mode, on which surface, with which tools, for which task, under which permission boundary, at what cost, with what evidence, and with whose approval. That question is slower than picking a model name, but it is the difference between a launch experiment and a dependable production practice.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Access Free Prompt Library →

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this