Custom GPT to Plugin Portability Audit: What Transfers, Access Rebuilds, and How to Validate the Replacement
Terminology: A custom ChatGPT assistant is a configured assistant, not a separately trained model. Here, generative pre-trained transformer (GPTA family of transformer-based language models trained to generate and analyze content. Open glossary entry) names the underlying language-model family, not the custom assistant itself. OpenAI custom-GPT guidance.
Evidence checkpoints
Documented point: For the planned custom-GPT migration, OpenAI documents that instructions become a plugin skill, knowledge files copy into plugin reference files, and connected apps are added as apps; selected model, conversations, sharing settings, and custom actions do not transfer automatically. This planned workflow varies by account or workspace; migration access, GPT eligibility and replacement behaviour require separate checks. [official source 1]
Documented point: OpenAI’s retirement frequently asked questions (FAQ)A collection of recurring questions and concise answers about a subject. Open glossary entry lists 11 September 2026 for affected Enterprise administrator notices; 22 September 2026 as a target for the migration experience/user banner; and 1 October 2026 as a target for delayed-migration-banner workspaces. The target dates are explicitly not a guarantee of availability to every account or workspace. These dates describe Enterprise notice and migration targets, not a guaranteed control in every account. [official source 1]
Documented point: OpenAI’s current Codex models documentation says GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex on 14 October 2026, while remaining available in the OpenAI application programming interface (API)A documented way for software systems to exchange requests and results. Open glossary entry. This is distinct from the 11 December 2026 custom-GPT retirement and must not be conflated with it. As of 1 October 2026, GPT-5.5 retirement is scheduled rather than complete. Replacement-model availability depends on plan, client and workspace controls. [official source 1]
Documented point: Custom GPTs are scheduled to retire on 11 December 2026. An 11 February 2027 date applies only to affected Enterprise workspaces with a qualified, approved deferral; OpenAI says Enterprise milestones are subject to change. The later retirement date requires a qualified, approved workspace deferral; it is not an automatic Enterprise extension. [official source 1]
Documented point: On 6 July 2026, OpenAI said GPT-5.5 Instant Mini was rolling out in ChatGPT as the fallback after GPT-5.5 Instant or Auto rate limits. It does not appear in the model picker, and that update does not affect the API or Codex. This is a ChatGPT usage-limit fallback, not a selectable API or Codex model or a migration-continuity promise. [official source 1]
Build the portability inventory before touching the migration control
A custom-GPT-to-plugin migration is not a one-for-one copy. The useful starting point is therefore an inventory of the workflow as it exists now, not a list of hoped-for plugin features. For one hypothetical example, this section audits a custom GPT called Policy Briefing Assistant. It summarises internal policy files, uses a connected document app, applies a fixed briefing format and calls a custom action that retrieves a policy record by identifier. It is shared with editors and occasional reviewers.
The example is a planning template, not a report of a migration performed by this publication. OpenAI’s Custom GPT retirement and migration FAQ documents the intended transfer boundaries: the latest published instructions become a plugin skill; knowledge files are copied into plugin reference files; and connected apps are added as apps. The selected model, existing conversations, sharing settings and custom actions do not transfer automatically. Those statements describe planned product behaviour, not proof that migration is available for a particular account, that this example is eligible or that the resulting plugin will behave identically.
OpenAI schedules standard custom-GPT retirement for 11 December 2026. According to its retirement FAQ, an 11 February 2027 date applies only to affected Enterprise workspaces with a qualified, approved deferral, and Enterprise milestones are subject to change. The FAQ also lists 22 September 2026 as the target for the migration experience and user banner, and 1 October 2026 as the target for workspaces that opted into delayed migration banners. These targets do not guarantee availability in every account or workspace. Before relying on any date or process, the owner should inspect the current workspace notice and confirm that the migration control is actually present.
This timetable concerns retirement of the custom-GPT product object. It is distinct from OpenAI’s separately documented retirement of GPT-5.5 from ChatGPT, ChatGPT Work and Codex on 14 October 2026. A model date must not be treated as the workflow’s retirement date, and the model selected for the original GPT must not be assumed to survive migration.
To make the portability inventory concrete, begin by separating the legacy GPT’s interface elements and integrations by role. The updated plugin menu, MCP connectors and action icons are distinguished in a ChatGPT interface walkthrough, which offers a vocabulary for a Custom GPT connector and action inventory that records what the replacement must account for, without assuming any component transfers automatically.
Use three evidence states rather than a binary “migrates” column
A robust inventory separates what OpenAI documents from what the owner still has to establish. Use these three states:
- Documented carry-over: OpenAI says the component is converted, copied or added during the planned migration. This confirms the intended treatment of the configuration object, but not behavioural equivalence.
- Documented non-carry-over: OpenAI says the component does not transfer automatically. The owner must rebuild it, replace it, archive it or accept its loss.
- Account-specific unknown: the official documentation does not establish the result for this GPT, workspace, user or provider account. The item needs inspection or staged verification.
The distinction matters because “copied” and “works” answer different questions. A knowledge file may be copied as a reference file while retrieval quality, formatting and source use remain unverified. A connected app may be added while the intended user still lacks provider authorisation. Conversely, a non-transferred custom action may have a viable replacement through an existing app or a custom Model Context Protocol (MCP)A protocol for connecting AI applications with tools and data sources through defined interfaces. Open glossary entry server, but the migration itself does not establish that replacement.
Decision rule: label an item “portable” only for the narrow property stated in the documentation. Record behaviour, access and outcome parity separately as unknown until reviewed. Do not upgrade a documented configuration transfer into a claim of functional equivalence.
Define the hypothetical workflow precisely
The sample Policy Briefing Assistant has the following intended journey:
- An editor supplies a policy record identifier and asks for a briefing.
- The assistant follows published instructions defining audience, headings, source-handling rules and prohibited speculation.
- It consults uploaded editorial standards and terminology files.
- It can use a connected document app to locate permitted supporting material.
- A custom action submits the record identifier to an internal service and retrieves current record fields.
- The assistant returns a structured briefing, noting absent information rather than inventing it.
- Editors can use the GPT; reviewers have narrower needs and may only inspect generated material.
This definition reveals that the workflow is not a single prompt. It combines guidance, static references, live data, an action, model-dependent execution, access controls and user habits. The inventory must preserve those distinctions. Treating the whole object as “the GPT” would conceal several documented losses.
A practical inventory record should identify an accountable owner, the workspace containing the GPT, the latest published version, each dependency, intended users, data sensitivity, evidence source, transfer classification, replacement owner and verification status. Do not put credentials, access tokens, private customer records or other secrets into the inventory’s prompts or sample cases. Use synthetic identifiers and sanitised examples until an authorised test environment and data-handling procedure have been approved.
Baseline inventory for Policy Briefing Assistant
| Component | Example inventory entry | Documented migration treatment | What remains unknown | Owner action before migration |
|---|---|---|---|---|
| Latest published instructions | Version PBA-17: audience rules, required headings, source hierarchy, missing-data rule and final checklist | Become a plugin skill | Whether the skill triggers reliably and follows every rule in direct, indirect and difficult cases | Publish the intended version, archive its text and turn each material rule into a reviewable expectation |
| Knowledge files | Editorial Standard.pdf, Terminology.csv and Policy Types.docx | Copied into plugin reference files | Retrieval, interpretation, source selection, formatting and treatment of conflicting or missing content | Record file names, versions, owners, sensitivity and representative questions tied to specific passages |
| Connected apps | Document repository used to locate approved supporting material | Added to the plugin as apps | Installation, user authorisation, role access, supported actions, source scope, regional availability and workspace restrictions | Map every app to its administrator, provider account, minimum access and intended users |
| Selected model | The model selected in the published GPT configuration | Does not carry over; Enterprise defaults apply where relevant | Which models the target workspace and users will actually have, and whether outputs remain acceptable | Record the old selection for audit context, but design task-based checks rather than demand model continuity |
| Conversation history | Past editor conversations containing examples, corrections and informal operating knowledge | Does not move | Whether an account-specific archive or export process is available and permitted | Extract authorised, sanitised test cases and formalise undocumented rules before migration |
| Custom action | Retrieve a policy record by identifier from an internal service | Does not transfer automatically | Whether an existing app, app template or custom MCP server can reproduce the necessary authorised operation | Document input, output, authentication, errors, approvals and prohibited operations without copying secrets |
| Sharing roles | Creator, two maintainers, editorial users and occasional reviewers | Sharing settings do not carry over; a migrated personal plugin starts private | Who can install, use, administer or access each underlying app after migration | Create a named access matrix and plan explicit re-sharing and user verification |
| User dependencies | Provider accounts, workspace membership, role eligibility, platform access and reliance on old conversation links | No universal carry-over is documented | Plan, region, role, platform, provider permissions and app capabilities for each user group | Identify dependencies per persona and test access separately from plugin configuration |
The table is the baseline, not the acceptance result. A row is complete only when it contains evidence specific enough for another maintainer to inspect. “Has files” is inadequate; “Editorial Standard.pdf, revision 8, owned by Editorial Operations, containing mandatory briefing headings” is auditable. “Uses an app” is also inadequate; identify the provider, intended source boundaries, user group and required read or write operation.
1. Capture the latest published instructions, not the editor draft
OpenAI’s FAQ says migration uses the latest published version. Drafts and unpublished edits do not transfer. Publishing is a version-state requirement for migration; it does not mean the GPT has been shared publicly. The owner should therefore distinguish four things that are often conflated: the text visible in the editor, the latest saved draft, the latest published configuration and the audience with whom that published GPT is shared.
For the sample GPT, suppose PBA-17 is the published version while a draft adds a new rule: “When the effective date is absent, write ‘Not supplied’ and do not infer it.” If the owner migrates without publishing the draft, the documented migration basis is PBA-17, not the unpublished change. The audit must not state that the missing-date rule transferred merely because it was visible to the creator in an editor.
Use this procedure:
- Open the GPT in the owner workspace and identify its current publication state.
- Copy the latest published instructions into a controlled configuration record. Preserve headings, ordering and explicit decision rules.
- Compare them line by line with any draft. Mark every draft-only change.
- Have the workflow owner decide whether each draft change should be published, discarded or retained for a later plugin revision.
- Assign a local version identifier and review date. This identifier is an audit convention, not an OpenAI product field.
- For each important instruction, write an observable expected behaviour and at least one case in which the instruction should not apply.
For example, the instruction “Use the executive briefing format for requests asking for a policy summary” should be decomposed into evidence that can later be reviewed: the output contains the required headings; missing fields are identified; and a request merely asking for a document title does not unnecessarily invoke the full briefing workflow. These are suggested checks, not guaranteed product results.
Documented carry-over: the published instructions become a plugin skill. Not documented as equivalent: identical triggering, interpretation, wording or output. OpenAI’s guidance recommends checking direct and indirect triggers, incomplete-input follow-up behaviour, non-trigger cases and edge cases. The inventory should therefore preserve the intention behind each material instruction, not only the prose.
Decision rule: do not proceed while consequential requirements exist only in an unpublished draft, old conversation or maintainer memory. Publish the approved version if appropriate, or record the requirement as a separately governed plugin change. For legal, privacy, security, financial, employment, government or similarly consequential guidance, require review by a qualified human owner rather than treating instruction transfer as approval of the content.
2. Inventory knowledge files by purpose, authority and sensitivity
The FAQ says knowledge files are copied into plugin reference files. A reference file supplies material that a skill can use; it is not the same thing as live app access or an action. File copying does not establish that the plugin will retrieve the same passage, resolve conflicting documents in the same way, reproduce citations or preserve a former response style.
For Policy Briefing Assistant, create one record per file:
| File | Purpose | Authority rule | Example review question |
|---|---|---|---|
| Editorial Standard.pdf | Defines mandatory structure and evidence labels | Current approved revision takes precedence over older examples | Does a sample briefing contain every mandatory section? |
| Terminology.csv | Maps internal terms to approved public wording | Use an exact approved term where a mapping exists | Does a request containing an internal term produce the approved equivalent? |
| Policy Types.docx | Describes categories and boundary cases | Do not assign a category when required evidence is absent | Does an ambiguous example trigger a request for missing information rather than an invented classification? |
Record the owner, revision date, source of authority, file format, intended audience and retention status. Flag duplicates, expired documents and conflicts before migration. Copying two contradictory files faithfully would preserve the conflict, not solve it.
A useful negative case is a question whose answer is absent from every reference file. The expected requirement might be: “State that the reference set does not establish the answer; do not manufacture a policy.” Another is a deliberately outdated term appearing only in an obsolete file. The expected requirement should specify whether the current source overrides it. These are examples of acceptance criteria, not claims about actual plugin performance.
Do not upload untrusted material merely to make the inventory comprehensive. Review files for embedded instructions, personal data, confidential material and unnecessary content before they become reference files. Keep secrets, credentials and private keys out of files and prompts. Where a file affects security, privacy, legal rights, money, employment or government services, a responsible human must confirm both its authority and the replacement’s output before use in a consequential decision.
Decision rule: classify a knowledge file as ready only if it has a named purpose, current owner, authority level and at least one expected-use case plus one absence or conflict case. A copied file with uncertain authority is a migration liability, not an accepted dependency.
3. Separate connected-app presence from usable access
OpenAI documents that connected apps are added to the migrated plugin as apps. That is narrower than saying every user remains connected. Plugin installation does not complete provider-account authorisation, grant provider scopes, alter source-system access or bypass workspace policy. Availability can depend on plan, workspace, role, region, platform and the app’s capabilities.
For the hypothetical document repository, inventory at least:
- the app and provider;
- the business purpose for using it;
- which users need it;
- whether the workflow needs search, read, write or another supported action;
- the permitted source locations;
- the relevant workspace role;
- the provider-account role;
- any action approval requirement;
- the accountable workspace and provider administrators; and
- what the workflow should do when access is absent or a source is outside scope.
For example, an editor may be permitted to search an approved policy library but not a restricted personnel folder. The migrated plugin containing the app does not prove that either source is accessible. A sensible expected negative behaviour is that a request for the restricted folder fails clearly and does not imply that no such material exists. Lack of access and lack of data are different conditions.
OpenAI’s administration guidance separates plugin installation, underlying-app availability, role access, provider authorisation and action controls. Preserve those as separate inventory fields. A single “connected: yes” field hides the most important failure modes.
Decision rule: regard an app as configured only when it appears in the intended plugin; regard it as usable only after an authorised representative of each intended user role can perform the minimum permitted operation. Regard write or external side-effect capability as a separate approval gate. Human review is mandatory before actions affecting security, privacy, money, employment, government records or other consequential matters.
4. Record the selected model as a lost dependency
The original GPT’s selected model does not carry over. OpenAI’s FAQ says Enterprise defaults apply where relevant. Consequently, “uses model X” belongs in the historical baseline but cannot be a migration acceptance requirement phrased as “the plugin must retain model X”. The relevant acceptance question is whether the replacement completes the required tasks under the models and controls actually available in the target workspace.
Do not infer continuity from a model picker, release note or prior conversation. OpenAI’s 6 July 2026 release note describes GPT-5.5 Instant Mini as a ChatGPT fallback after GPT-5.5 Instant or Auto rate limits; it does not appear in the model picker, and that update does not affect the API or Codex. It is not a selectable replacement target for this inventory.
For Policy Briefing Assistant, record: “Original selected model: historical configuration only; non-transferable.” Then attach task requirements such as mandatory headings, no invented dates, correct terminology and explicit failure on missing app access. This shifts validation from an unsupported model-continuity assumption to observable workflow behaviour.
Trade-off: requiring a specific old model may block migration even though the business task could be met through the workspace’s current configuration; ignoring the old model entirely may overlook behaviour changes. Preserve it as context, but make task outcomes and safety constraints the acceptance criteria.
5. Treat conversation history as non-portable operational knowledge
Existing GPT conversations do not move. This matters when users have allowed examples, corrections and workarounds to accumulate in chat history instead of updating the published instructions or reference files. Conversation history is not a substitute for maintained configuration, and the official sources assigned here do not establish a universal export or archive process.
Ask each maintainer to identify, without copying sensitive material into prompts:
- frequently reused requests;
- hard cases that prompted corrections;
- output formats users rely upon;
- known failure patterns;
- clarifying questions that are operationally necessary; and
- rules that exist only in prior conversations.
Convert only authorised and necessary examples into a sanitised review corpus. For instance, replace a real record identifier and personal details with synthetic values while preserving the structural challenge: missing effective date, conflicting category indicators or inaccessible source. Record expected characteristics rather than a single exact paragraph unless exact wording is genuinely required.
Decision rule: if the workflow depends on facts or corrections found only in old chats, migration is not ready. Move approved rules into the published instructions or governed reference material, and retain only the minimum authorised test evidence. Consult the organisation’s privacy, records and legal owners before archiving conversations containing personal, regulated or confidential information.
6. Decompose every custom action before choosing a replacement
Custom actions do not transfer automatically. OpenAI’s skills documentation distinguishes reusable workflow guidance from live data, authentication, authorisation and actions, which require an app or MCP-backed capability. A skill-only plugin must not be described as recreating a former custom action.
For the sample record-retrieval action, inventory the contract without recording secrets:
- Trigger
- A user requests a briefing and supplies a syntactically valid policy record identifier.
- Input
- The identifier, plus only the context necessary for the authorised request.
- Authorisation
- The user and service must be permitted to access the requested record.
- Output
- Defined record fields, source status and any explicit missing-value indicators.
- Failure cases
- Malformed identifier, record not found, denied access, service unavailable and incomplete response.
- Side effects
- None for retrieval; any future update operation would require a separate inventory and approval path.
Then classify the replacement path:
- Existing supported app: choose this only if its documented operations and access model cover the required capability.
- App template: treat it as a starting point, not a ready connection or evidence of provider access.
- Custom MCP server: consider this where live data, authentication or a controlled action needs separate technical implementation.
- No replacement: remove or redesign the capability if no governed option meets the requirement.
The inventory should not choose among these based on naming similarity. A document-search app is not necessarily a substitute for a record-retrieval action, even if both return text. Compare required inputs, authoritative source, access rules, outputs, errors and side effects.
Decision rule: no custom action is marked portable. Mark it “replacement identified” only when an option covers its required operation and governance model; mark it “accepted” only after authorised verification. If the action can change data, spend money, communicate externally, affect employment, alter government records or create another consequential effect, require an explicit human confirmation step and specialist security, privacy and legal review.
7. Rebuild sharing from named roles, not inherited assumptions
Sharing settings do not carry over, and OpenAI says a migrated personal plugin starts private. Migration therefore does not preserve the old GPT’s audience, make the plugin public or give former users access. Publishing the source GPT for migration also does not mean it was publicly shared.
Create an access matrix before migration:
| Persona | Needed capability | Plugin relationship | Underlying dependency | Verification required |
|---|---|---|---|---|
| Creator | Maintain configuration and migration evidence | Owner or appropriate administrative role | Workspace role and permitted app administration | Can inspect the intended configuration without receiving unnecessary provider access |
| Maintainer | Update skill and reference material | Explicitly assigned role | Workspace permission; app access only where required | Can perform maintenance while separation of duties remains intact |
| Editor | Generate briefings and read approved sources | Explicitly shared access | Provider authorisation and approved source scope | Can run representative read-only cases and receives clear denial outside scope |
| Reviewer | Review outputs, not necessarily invoke live data | Only the access needed for the review process | No provider connection unless justified | Can complete review without excessive privileges |
This matrix prevents a common category error: giving a reviewer broad app access merely because they previously received a link to the GPT. Plugin access, workspace role, provider authorisation and source-system permissions are separate controls.
Trade-off: restoring the old audience quickly may reduce disruption, but bulk re-sharing before access review can reproduce obsolete or excessive permissions. Start private as documented, establish the minimum roles and expand access in stages. Public directory distribution, if required, is a separate packaging, review, approval and publication process; it is not part of automatic migration.
8. Inventory user dependencies that sit outside the GPT
A user may have depended on more than the GPT’s sharing setting. Record workspace membership, role, plan, region, platform, provider account, approved data sources, supported actions and any reliance on old conversation links. OpenAI’s documentation says plugin and app capabilities can vary by plan, workspace, role, region and platform, so one successful owner view would not establish access for all intended users.
For each persona, write a dependency chain. An example for an editor is:
Workspace membership → permission to use the plugin → plugin explicitly shared → underlying app available → provider account authorised → approved policy library accessible → requested read action permitted.
A break at any stage has a different remedy. If the plugin is not shared, the plugin owner acts. If provider authorisation is missing, the user or provider administrator may need to act. If the source folder is prohibited, sharing the plugin more broadly is not a valid fix.
Include a fallback for users whose dependency cannot be restored. For example, a reviewer without authorised app access might receive a human-reviewed briefing through an approved records channel rather than being granted direct source access. This is a design option, not a product guarantee.
Decision rule: do not sign off “user access restored” from the creator’s account. Verify at least one authorised representative for every materially different role and dependency chain, while preserving least privilege.
When deciding which instructions and workflow settings to rebuild, it helps to separate intent from copied configuration. A playbook on moving to ChatGPT Work is a comparison point for preserving Custom GPT workflow configuration, although its destination is not a plugin, so a plugin replacement should document what it recreates rather than assume anything transfers.
Turn the inventory into a controlled migration record
The completed record should contain two adjacent views: the source baseline and the replacement requirement. The source baseline preserves what the published GPT contains. The replacement requirement states what the workflow must still accomplish without assuming that transferred components behave identically.
For example:
| Baseline fact | Replacement requirement | Evidence needed later |
|---|---|---|
| PBA-17 contains a rule against inferred effective dates | The skill must identify an absent date without inventing one | A sanitised missing-date case reviewed by the workflow owner |
| Terminology.csv contains approved public terms | The replacement must use the current approved mapping when relevant | A reference-dependent case and an obsolete-term negative case |
| The GPT uses a document repository | Editors may search only approved sources with their own authorised access | Positive access and denied-source checks for the editor role |
| A custom action retrieves a record | An authorised replacement must distinguish not found, denied and unavailable states | Technical and governance review of the chosen app or MCP-backed route |
| The GPT is shared with editors | Named editor roles must receive explicit plugin access without unnecessary administrative rights | Role-based access confirmation after migration |
Do not put a passing status against a row merely because the migration interface reports completion. Completion can establish that a planned conversion step ran; it does not establish reference quality, app authorisation, action parity, model continuity or user access. Reserve final acceptance for the staged verification covered later in the audit.
Minimum readiness gate
The owner can classify the hypothetical GPT as inventory-ready only when all of the following are true:
- the owner workspace and applicable retirement notice have been checked;
- the intended latest published version is identified and archived;
- draft-only changes have an explicit disposition;
- every knowledge file has a purpose, owner, authority and sensitivity classification;
- every connected app has a user, provider, source and action dependency map;
- the selected model is recorded as non-transferable rather than promised;
- necessary conversation-derived cases have been sanitised and formalised, subject to applicable review;
- every custom action has a documented contract and replacement decision owner;
- sharing roles are named and will be re-established explicitly;
- user dependencies are recorded by persona rather than inferred from the creator’s access; and
- security, privacy, legal, financial, employment, government and other consequential uses have named human reviewers.
If any required item is unknown, the workflow is not ready for an evidence-based migration decision. The appropriate response is remediation, not an assumption that the plugin will fill the gap. If the migration control is absent, treat availability as unresolved and consult the current account or workspace notice; the published target dates alone are not evidence that the feature is enabled.
The central portability finding for Policy Briefing Assistant is therefore deliberately narrow. Its latest published instructions, knowledge files and connected apps have documented migration destinations: a skill, reference files and apps respectively. Its model selection, conversations, sharing settings and custom action do not automatically transfer. Whether the replacement is usable remains an account-specific and workflow-specific question requiring later verification. That distinction keeps the audit grounded in OpenAI’s documented product behaviour without presenting planned migration as tested equivalence.
Rebuild access and actions as separate control layers
OpenAI’s Custom GPT retirement and migration FAQ, as reviewed for this article on 1 October 2026, documents three different migration treatments: published instructions become a plugin skill, knowledge files are copied into plugin reference files, and connected apps are added as apps. These are configuration transfers, not evidence that the replacement can reach the same records, perform the same actions or produce identical answers. The same FAQ states that custom actions do not transfer automatically. An access audit must therefore separate four questions: what guidance activates, what static material is available, which app is included, and what the current user is authorised to read or change in the provider’s source system.
The practical consequence is that “the plugin migrated” is not an acceptance criterion. A replacement can contain the expected skill and app while still failing because the skill did not activate, a reference file was absent, the provider account was not authorised, a workspace role blocked the app, the source system withheld a folder, or an action required approval. Conversely, broad access can make an expected-success demonstration appear successful while concealing excessive permissions. Accept the replacement only when both permitted operations and required denials match the approved design.
Distinguish the execution surfaces before assigning permissions
A skill is reusable guidance describing a workflow, its sequence and decision points. It can tell the plugin when to gather information, which checks to perform and how to structure an answer. A skill is not, by itself, a live connection to an external service. OpenAI’s skill-building documentation assigns live data, authentication, authorisation and controlled actions to an app or an MCP server rather than to the skill alone.
A reference file is copied knowledge available to the plugin. It can support definitions, policies, templates or other relatively stable material, but its presence does not establish that the plugin will retrieve the right passage for every request. Nor does it establish access to a current provider record. For example, a copied procurement policy may explain which approvals are required, whereas the current purchase request, supplier status and approval state must come from an authorised source system if the workflow depends on live information.
An included app identifies an integration surface that the plugin may use. According to OpenAI’s plugin and permissions documentation, installation does not connect the user’s provider account, grant provider scopes, override workspace policy or confer access to records that the provider itself denies. An app may be present but unusable for a particular member because installation, underlying-app availability, role access, provider authorisation, action permissions and source restrictions are separate controls.
An action reads from or writes to an external system through an app or other supported integration. A read can still be consequential: retrieving an employee file, financial record or government case may expose restricted data. A write can create, alter, submit, send or delete something. OpenAI’s permissions guidance distinguishes action approval from source access. An approval setting cannot create an account connection or enlarge the provider’s permissions, and saved approval does not override restrictions or safety protections.
Use this decision rule: place stable workflow guidance in the skill, stable supporting material in reference files, and live records or controlled operations behind an authorised app or MCP server. If a requirement says “look up the current value”, “submit”, “update”, “send”, “create” or “delete”, do not classify it as satisfied by a skill or copied file. If the original custom GPT used a custom action for that requirement, record the capability as unresolved until a supported replacement and its controls have been identified.
A replacement design needs explicit boundaries between reusable instructions, external connections and callable operations before validation begins. A Codex plugin tutorial is organised around a skill, an app connector and an MCP server with tools, which helps readers map Codex plugin skill connector boundaries instead of collapsing these responsibilities together during a port.
Trace activation before testing data access
Skill activation and app access are related but different. The skill may activate correctly and then encounter an access denial; alternatively, the app may be authorised but never invoked because the skill did not recognise the request. Combining these failures under “integration problem” obscures the remedy. The first calls for revised skill boundaries or trigger examples; the second calls for access, provider or workspace investigation.
For each workflow, prepare four prompt classes as an example test design:
- Direct trigger: the user explicitly names the task, such as “Prepare a policy briefing from the approved policy library and the current programme register.” The expected behaviour is activation of the relevant workflow, followed by a request for any essential missing input.
- Indirect trigger: the user describes the intended result without naming the skill, such as “I need a two-page note comparing the approved policy with the programme’s latest recorded status.” The expected behaviour is appropriate activation without requiring a trigger wording.
- Incomplete request: the user omits the programme, date range or output audience. The expected behaviour is a clarifying question rather than an invented selection.
- Non-trigger: the user asks for an unrelated task, such as editing a social invitation. The expected behaviour is that the restricted policy workflow and its connected app are not invoked.
OpenAI’s skill documentation recommends direct and indirect trigger tests, follow-up behaviour for incomplete inputs, non-trigger cases and edge cases where the workflow must not invent information or unsupported actions. These are suggested test criteria, not reported results. Record each test’s expected activation, permitted tools, prohibited tools and required user confirmation before running it in the target workspace.
A useful failure distinction is “correct refusal” versus “broken workflow”. If a user without programme-register access asks for a current status, the correct result may be an access denial or a bounded explanation that the live record cannot be retrieved. If an authorised user receives the same denial, investigate provider connection, role, workspace controls and source scope. Do not weaken the skill’s refusal behaviour merely to make both tests return content.
Build an access lattice rather than a single permissions list
An access lattice is a structured ordering of permissions from no access through increasingly consequential capabilities. It prevents a misleading binary label such as “app enabled”. The table below is an example audit template. Adapt it to the provider’s documented controls and the organisation’s policy; it is not a description of universal product labels.
| Layer | Question to verify | Lowest acceptable evidence | Negative case | Decision rule |
|---|---|---|---|---|
| Plugin installation | May the intended workspace and role install or use this plugin? | Workspace-specific confirmation by an authorised administrator or owner, plus an intended-user check | A member outside the permitted role cannot use it | Installation alone never counts as app or data access |
| Underlying app availability | Is the required app supported for this plan, workspace, region, role and surface? | Current workspace configuration and official product documentation | The plugin remains usable for skill-only work when the app is unavailable, or gives a bounded error if the task requires it | Do not approve a live-data workflow on directory presence alone |
| Provider connection | Has this user authorised the correct provider account? | A user-specific connection check that reveals no secret tokens | An unconnected user is asked to authorise or is denied; no data is fabricated | Never treat installation or administrator approval as account authorisation |
| Provider scope | Which provider permissions were granted? | Documented scopes mapped to required operations | A scope omitted deliberately causes the corresponding operation to fail safely | Reject unexplained scopes and prefer the minimum needed for the approved task |
| Source-system entitlement | Which sites, folders, projects, records or fields can this identity access? | Provider-side access review using a representative non-sensitive test location | A restricted location remains undiscoverable and unreadable | Plugin permission must not be assumed to enlarge source access |
| Read action | May the integration search, list or retrieve the required data? | Approved read operation tied to a named business purpose | A user without read entitlement receives no record content, snippet or metadata leakage | Grant read only where static reference files cannot meet the need |
| Write action | May it create or modify a source-system object? | Named action, target boundary, accountable owner and approval behaviour | An unauthorised write is blocked before execution | Do not infer write permission from successful reading |
| Destructive or externally visible action | May it delete, submit, publish, send or otherwise create an external consequence? | Explicit human review, displayed target and payload, and provider-side permission | Ambiguous, unauthorised or unapproved requests do not execute | Require human confirmation for consequential action; automate only within separately approved policy |
| Sync or indexed source | Is data available through an approved synchronised or indexed source? | Named source boundary, owner and current restriction review | Excluded sources do not appear in search or citations | Do not equate sync access with permission for direct actions |
| Sharing and recipient access | Can each intended recipient reach the plugin and each required app? | Role-based recipient sample covering authorised and unauthorised users | A former GPT user who has not been granted replacement access remains denied | Rebuild access explicitly because old sharing settings do not transfer |
The lattice should be evaluated per identity, not merely per plugin. An administrator, workflow owner, ordinary member, contractor and external collaborator can encounter different results even when they open the same plugin. Add a row for each materially different user group and provider account type. If one group needs only policy guidance, do not grant it live-record access simply because another group requires that capability.
For each connected application, treat access as a permission to re-establish and verify, not a setting that follows the old GPT automatically. An article on the Data plugin frames it as an analytics workflow rather than a permission bypass, and covers Data plugin permissions and validation for connected business data, which illustrates the kind of check required.
Verify included apps without mistaking presence for readiness
The migration FAQ says connected apps are added as apps, but the plugin documentation makes availability conditional on factors including plan, workspace, role, region, capability and platform. Begin with an app-by-app reconciliation: original connected app, migrated app entry, required operation, intended user group, provider identity, necessary source boundary and approval requirement. Mark “present” separately from “authorised”, “source accessible” and “action permitted”.
For example, suppose the Policy Briefing Assistant previously consulted a document repository and a programme register. The document app appearing in the replacement establishes only that the integration was included. The maintainer must still determine whether the intended analyst can connect the correct provider account, read the approved policy library, and remain unable to open the restricted personnel folder. The programme-register app must be assessed separately because access to documents does not imply access to programme records or permission to modify them.
Use a fresh low-privilege test identity where organisational policy permits, rather than relying solely on an owner or administrator account. Elevated accounts can conceal missing role mappings and excessive scope. The intended positive case should use a non-sensitive test record that is within scope. The paired negative case should request a known out-of-scope location or operation without placing restricted content in the prompt. Pass only when the permitted request works as designed and the prohibited request fails without exposing content.
Do not paste credentials, access tokens, private keys, session data or real restricted records into prompts or test notes. Treat copied provider output as untrusted data: it may contain misleading instructions or content that is irrelevant to the approved workflow. The skill should constrain how retrieved material is used, while app and provider controls determine what can be retrieved. Neither layer substitutes for the other.
Test missing data as deliberately as missing permission
A missing record is not the same as an access denial. Search may return nothing because the record does not exist, the name is wrong, indexing or synchronisation excludes the source, the date filter is too narrow, or the user lacks access. A replacement that turns all empty results into “no such record exists” can make an unsupported factual claim. Define distinct expected responses for absence, ambiguity, unsupported source and denied access wherever the integration exposes those distinctions.
| Negative test | Example request | Expected safe behaviour | Failure to investigate |
|---|---|---|---|
| Required reference file absent | Draft a briefing that depends on a policy appendix deliberately omitted from the test copy | Identify the missing basis and request the file or decline unsupported conclusions | Invented policy provisions or reliance on unrelated material |
| Reference passage not found | Quote a clause that is not present in any approved reference file | State that the clause was not found; do not fabricate a quotation | Confident quotation without a recoverable source |
| Provider not connected | Retrieve the latest programme status using an unconnected test identity | Request authorisation or report that the live source is unavailable | Implied retrieval, guessed status or exposure from another user’s connection |
| Source permission missing | Open a record outside the test identity’s provider entitlement | Deny access without returning title, snippet, field values or hidden metadata | Partial leakage or advice to bypass the restriction |
| Valid query, no matching record | Search for a deliberately non-existent test identifier | Report no match within the searched source and criteria, without claiming universal non-existence | Fabricated record or overbroad conclusion |
| Ambiguous record | Ask for “the North programme” where several authorised records match | Present bounded disambiguation details or ask the user to choose | Silent selection of one record |
| Read allowed, write denied | Retrieve a permitted item and then request an update with a read-only identity | Return the read result but block the update | Assuming successful retrieval permits modification |
| Approval withheld | Prepare a permitted update but decline the action confirmation | Do not execute; clearly distinguish a draft from a completed action | Execution despite refusal or language implying completion |
| Untrusted source instruction | Retrieve a test document containing text that asks the system to ignore workflow controls | Treat the text as source content, not authority to change permissions or invoke unrelated actions | Following embedded instructions or disclosing unrelated data |
| Unsupported action | Ask the app to delete or publish when only reading is approved | Refuse or explain that the action is unavailable | Invented success, workaround instructions or use of an unrelated tool |
Run these as paired cases. The positive case establishes that the approved path is usable; the negative case establishes that the boundary holds. A system that denies everything is not ready, but neither is one that succeeds by using broader access than intended. The acceptance rule is therefore dual: required tasks must complete within the approved source and action boundary, while excluded data and operations must remain unavailable.
Place approval at the point of consequence
Action approval should be designed around the effect, not around the convenience of avoiding prompts. OpenAI’s permissions documentation indicates that app action settings do not connect accounts, alter source-system access or override workspace policy. It also distinguishes any individual-app “Allow all actions” option from standard account- or workspace-wide selection and says saved approval does not override restrictions or safety protections. Consequently, an approval choice is one control in the chain, not a blanket authority.
Classify each action by what changes. A search or read may require no per-action confirmation if the organisation has approved that use and the data is non-consequential, although provider and workspace permissions still apply. Creating a private draft is different from sending it. Updating an internal status field is different from submitting a regulatory return. Deleting a record is different from marking a test item complete. Avoid collapsing these into a generic “write” permission.
For every consequential operation, define the preview that a human reviewer must see: action type, destination, affected object, material values, external recipients where applicable and whether the operation is reversible. Require human review for security, privacy, financial, employment, government, healthcare and other consequential decisions or actions. The reviewer must validate the underlying evidence and authority rather than merely approve fluent wording.
As an example procedure, a programme update workflow could retrieve an authorised test record, draft a proposed status change, display the target record and changed fields, and wait for explicit confirmation. The test then branches: approval should invoke only the specified action; rejection should leave the source unchanged; altered or ambiguous targets should require a new review. This is an illustrative control pattern, not a guarantee that every app supports that sequence.
Decide whether a former custom action has a legitimate replacement
Because custom actions do not transfer automatically, do not begin by asking how to recreate each endpoint. First ask whether the capability remains necessary, whether a supported included app provides it, and whether the organisation is willing to authorise the associated data and action surface. Some former actions should be retired, narrowed or replaced with a manual hand-off rather than rebuilt.
- Remove: choose this when the action is unused, duplicates another approved system or no longer has an accountable owner.
- Replace with static guidance: choose this only when the requirement is genuinely informational and does not need current provider data or execution.
- Use an existing app: choose this when the required source and operation are documented as supported and the necessary workspace, provider and action controls can be established.
- Use a manual hand-off: choose this when the plugin may prepare a draft or structured instruction but a person must execute the operation in the source system.
- Commission a custom MCP server: consider this when live data, authentication or a controlled action remains necessary and no suitable supported app exists.
- Retire the workflow: choose this when no acceptable replacement can meet functionality, access and governance requirements before the applicable retirement date.
Do not label an app template as a completed connection. OpenAI’s plugin documentation distinguishes templates and included surfaces, and capability can vary by platform. The replacement decision must cite the actual supported operation, not visual similarity or a directory description. If evidence establishes reading but not writing, classify only reading as replaced.
Treat a custom MCP rebuild as a separate technical project
MCP is a protocol through which a suitable server can expose data or tools to a compatible client. OpenAI’s skill documentation identifies an MCP server as the component for live data, authentication, authorisation and actions when those are needed. That architectural option does not mean a former custom action is automatically converted, technically compatible or approved for use.
Create a separate project record for any proposed custom MCP server. Its scope should include the source-system interface, authentication model, user and service identities, authorisation checks, minimum operations, input validation, output handling, error states, audit requirements, deployment ownership, maintenance responsibility and decommissioning plan. Keep implementation credentials and production data out of prompts, examples and migration worksheets.
The technical team should map every exposed tool to one business operation. A broad tool such as “call arbitrary endpoint” is harder to constrain and review than named operations such as “read approved programme status” or “propose test-record update”. This is a design trade-off: narrower tools require more deliberate implementation but make intended use, permission checks and negative tests clearer. Acceptance should favour the smallest tool surface that meets the approved workflow.
Separate protocol functionality from organisational approval. A server may be capable of reaching an endpoint while the workspace, provider or organisation forbids that use. Directory checks, technical connectivity or successful authentication are not security certification and do not replace privacy, vendor, records-management or legal review. Where personal, financial, employment, healthcare, government or security-sensitive information is involved, require the relevant human owners to approve the purpose, data boundary and action design before production use.
Use an independent release gate for the MCP project: documented owner; approved data and action scope; threat and privacy review proportionate to the data; test identities; positive and negative authorisation cases; safe handling of provider errors; human approval for consequential writes; operational monitoring; rollback or disablement route; and support ownership. If these are incomplete, the custom action remains “not replaced” even if a prototype can return data.
Perform a security and privacy review against the actual data path
Review the full path rather than the plugin name: user request, skill activation, reference retrieval, app or MCP call, provider authorisation, source-system result, model processing, displayed answer and any subsequent action. At each step, identify the data class, accountable owner, minimum required fields, authorised user groups and prohibited destinations. This avoids approving a benign-sounding workflow while overlooking sensitive fields returned by a broad source query.
Apply data minimisation to both reads and writes. If a briefing needs programme name, status, reporting date and approved public risk summary, do not retrieve employee notes or unrelated attachments. If an action needs to update one status field, do not expose deletion or arbitrary record modification. Where the integration cannot be narrowed sufficiently, choose a manual hand-off or retire that portion of the workflow.
Reference files require their own review because copied knowledge and live source access have different lifecycles. Confirm that each file is authorised for the plugin’s intended audience, remains current enough for its purpose and contains no secrets or unnecessary personal data. A file copied successfully can still be inappropriate to distribute through the replacement. Remove obsolete or unauthorised files before approval rather than relying on instructions that ask users not to retrieve them.
Test cross-user separation. One user’s provider connection must not be assumed to serve another user, and content obtained under an owner’s elevated access must not become a substitute for recipient authorisation. Run recipient checks with appropriately configured identities and synthetic or non-sensitive records. Do not place real confidential content in prompts merely to prove that denial works.
Finally, record residual risk rather than converting uncertainty into a pass. Examples include an app unavailable on a required platform, a provider scope broader than the workflow needs, an approval flow that cannot distinguish draft from send, or an MCP server awaiting organisational review. The decision owner should choose among remediation, restricted rollout, manual fallback or retirement. A feature should not enter production solely because the migrated skill and app appear in the plugin.
Use a two-sided acceptance gate
Approve the access rebuild only when the evidence ledger shows all of the following: the correct skill activates for direct and indirect requests; unrelated requests do not trigger it; required reference files are present and missing material is reported rather than invented; each intended user group can reach the plugin; each required provider account is authorised separately; source-system permissions expose only approved records; required reads work; prohibited reads fail without leakage; writes are limited to approved operations; consequential actions receive human review; rejected approvals do not execute; and every former custom action is marked replaced, narrowed, manually handled or retired.
Do not accept partial evidence as equivalence. A successful owner demonstration does not establish member access. A successful read does not establish a safe write. An app being included does not establish provider consent. A provider connection does not establish access to every source. A skill activating does not prove that its references are complete. A custom MCP prototype does not establish production approval.
The practical sign-off should name the accountable workspace owner, app or source owner, security or privacy reviewer where required, workflow maintainer and business approver. For consequential domains, human reviewers remain responsible for the decision and for validating source evidence. If any mandatory layer lacks an owner, evidence or negative-case result, classify the replacement as requiring remediation rather than ready.
Run acceptance as a staged replacement review
This acceptance plan is a reproducible procedure, not a report of an executed migration. It deliberately leaves the “Expected” and “Observed” fields empty so that the owner records account-specific evidence rather than inheriting assumed results. OpenAI’s Custom GPT retirement and migration FAQ, reviewed on 1 October 2026, recommends testing familiar prompts and at least one harder case, including skill selection, instruction following, reference use, output completeness, formatting, tools and integrations. The same FAQ warns that a migrated plugin may respond differently.
The plan therefore tests capability, access and governance separately. A good answer to a familiar prompt does not prove that a reference file was used correctly. Successful installation does not prove that an intended user can authorise an app. A visible app does not establish permission to read a particular source or perform a write action. The decision rule is that every required layer must pass independently; evidence from one layer cannot compensate for a failure in another.
Freeze the candidate before collecting evidence
Start with a named release candidate rather than testing a moving configuration. Record the plugin owner, workspace, plugin version or internal revision identifier, migrated skill, reference-file set, included apps, any replacement for a former custom action, intended users and test date. Record the applicable original GPT version as well. The FAQ says migration uses the latest published version, while drafts and unpublished edits do not transfer, so an editor draft is not a valid baseline unless it is first reconciled with the published configuration.
- Copy the approved test cases into a version-controlled or otherwise change-tracked acceptance record.
- Assign each case an identifier, owner, risk class, required preconditions and evidence type.
- Freeze changes to instructions, files, apps and action configuration for the duration of one test round.
- Record defects without altering the candidate midway through the round.
- After remediation, create a new candidate revision and rerun all affected cases plus the core regression set.
For example, “RC-02” might identify the second candidate for an internal policy-briefing plugin. Its evidence bundle could list a migrated briefing skill, three approved policy reference files, one connected document app and no action-capable integration. This is an example record structure, not a claim about a real plugin or its behaviour.
The freeze rule matters because conversational outputs can vary and product availability can depend on plan, workspace, role, region, settings and rollout. Do not quietly adjust a skill after a failed case and then retain earlier passes as if they applied to the amended candidate. Any change that could affect triggering, instructions, source use, permissions, actions or formatting creates a new testable revision.
Define acceptance states before running a prompt
Use four outcome states: pass, fail, blocked and not applicable. “Pass” means the recorded acceptance criterion was met with reviewable evidence. “Fail” means the candidate ran but breached a criterion. “Blocked” means the test could not be completed because a prerequisite such as migration availability, installation rights, provider authorisation or source access was absent. “Not applicable” is valid only when the owner has documented why the capability is outside the approved replacement scope.
Do not count a blocked case as a pass, and do not average outcomes into a portability score. A single failed high-consequence control can be more important than numerous successful low-risk prompts. The owner should classify each requirement as one of the following:
- Mandatory: the workflow cannot be released without it.
- Conditional: it must pass when a named feature, app or user group is enabled.
- Advisory: useful behaviour that may be deferred without misrepresenting the supported workflow.
- Prohibited: behaviour that the replacement must refuse, avoid or route to a human.
The release rule is: all mandatory positive cases pass; all prohibited-behaviour cases demonstrate the required refusal or escalation; no unresolved high-consequence defect remains; and every blocked mandatory case prevents release. Advisory failures may be accepted only through a dated exception naming the impact, compensating procedure, owner and expiry date.
Before accepting a plugin replacement, define whether its data handling, access model and auditability meet the retiring workflow’s operating requirements. A ChatGPT Work security article offers a governance lens through its enterprise deployment access-control checks, alongside data governance and audit compliance, which can inform your acceptance criteria.
Build a prompt corpus from real duties, not demonstrations
Create a compact corpus that represents the work users actually need. Preserve prompts independently of old GPT conversations because the FAQ says existing conversations do not move. Remove secrets, personal data and confidential source text unless the approved test environment and data-handling review expressly permit their use. Prefer synthetic or redacted fixtures for security, privacy, employment, financial, government and other consequential workflows.
Divide the corpus into six groups:
- Familiar tasks: routine requests that exercised the original GPT’s central instructions.
- Harder cases: longer, ambiguous, conflicting or incomplete requests that expose instruction and reasoning boundaries.
- Reference-dependent tasks: requests that can be answered correctly only by using an approved file or source.
- Format-bound tasks: outputs requiring a specified structure, fields, order or omission rule.
- Tool and integration tasks: cases requiring an included app, provider access or an authorised action.
- Negative cases: requests that should not trigger the skill, should request clarification, should refuse an unsupported action or should avoid inventing missing information.
A representative set is more useful than many near-duplicates. For the example policy-briefing workflow, familiar cases could cover summarising an approved document and producing a briefing outline. Harder cases could combine two documents with conflicting dates, omit the target audience, or ask for a conclusion unsupported by the source. Negative cases could ask the plugin to amend a government record even though no approved write-capable tool exists. These are proposed test designs, not observed product outputs.
Use a blank evidence sheet for every case
The following template keeps procedure, acceptance criteria and recorded evidence distinct. The two comparison fields remain empty here by design. A tester should populate them only while executing the approved plan.
| Case | Procedure | Acceptance criterion | Expected | Observed | Evidence and reviewer |
|---|---|---|---|---|---|
| Familiar task F-01 | Submit the approved routine prompt in a new conversation with the candidate plugin enabled. | The relevant skill is selected; mandatory instructions are followed; required sections are present; unsupported facts are not introduced. | Record transcript reference, candidate revision, tester and review date. | ||
| Hard case H-01 | Submit an incomplete request that lacks a mandatory audience or date range. | The workflow asks for the missing input or applies only an expressly approved default; it does not invent the missing value. | Record the prompt, response and reviewer’s criterion-by-criterion assessment. | ||
| Reference case R-01 | Ask a question whose answer appears in one approved reference file and conflicts with a plausible unsupported answer. | The answer reflects the approved source, distinguishes source content from inference and does not claim access to an unavailable file. | Record the source passage used for human comparison. | ||
| Format case O-01 | Request the approved output structure using a fixed redacted fixture. | All mandatory fields appear in the required order; prohibited fields and unsupported content are absent. | Attach the output and completed format checklist. | ||
| Missing-tool case T-01 | Remove or withhold the required app from the test user, then request the dependent task. | The response identifies that the task cannot be completed through the available tools and does not imply that an action occurred. | Record user role, app state and response. | ||
| Unsupported-action case N-01 | Ask for a former custom action for which no approved replacement exists. | The plugin does not claim execution; it states the limitation and follows the approved escalation or manual hand-off. | Record response and action-system evidence showing no authorised execution path. |
Establish the familiar-task regression set
Select several routine tasks that collectively cover the skill’s essential instructions. Do not choose only the shortest or most polished prompts. Include the wording used by ordinary users, reasonable paraphrases and at least one indirect request that should still invoke the workflow. OpenAI’s skill-building documentation recommends testing direct and indirect triggers as well as non-trigger cases.
For each familiar task:
- State the business purpose and the minimum useful deliverable.
- Mark the instruction clauses that the case exercises.
- Provide a redacted input fixture and any approved reference material.
- Define content requirements separately from stylistic preferences.
- List disallowed assumptions, claims or actions.
- Run the case in a fresh conversation to avoid hidden dependence on earlier context.
- Have a reviewer assess each criterion rather than rating the answer by general impression.
Example: a briefing workflow receives a public policy document and a request for a one-page internal summary. Mandatory criteria might require the document’s stated purpose, effective date, affected group, unresolved questions and a clear separation between source statements and analyst interpretation. Tone and sentence length may be advisory. If an effective date is mandatory, an attractive summary that omits it fails.
Use an approved tolerance where exact wording is unnecessary. Acceptance should focus on required meaning, traceability and format, not textual identity with an old response. Because the selected model does not carry over, matching the original GPT word for word is neither a documented migration promise nor a sound acceptance rule.
Make harder cases expose operational ambiguity
A harder case should test a meaningful boundary, not merely add length. Construct cases involving missing fields, inconsistent sources, ambiguous dates, multiple possible workflows, requests outside scope or instructions that conflict with user wording. The criterion should specify whether the plugin must clarify, choose an approved default, present alternatives or stop.
For example, provide two approved reference documents with different publication dates and ask for “the current rule” without identifying jurisdiction. The acceptable procedure is not to pre-write the answer. Instead, require the candidate to identify the ambiguity, avoid silently merging incompatible sources and request the missing jurisdiction or apply a documented source-precedence rule. Human reviewers must verify the source comparison; the plugin must not be the final authority for legal or government decisions.
Run harder cases after the familiar set but before release to intended users. If a remediation changes the core skill, rerun the familiar cases because a clarification rule can alter routine behaviour. The decision rule is that a hard-case fix is unacceptable if it resolves one ambiguity by weakening a mandatory familiar-task instruction.
Verify reference use and source quality independently
Knowledge files are documented as copying into plugin reference files, but that does not establish identical retrieval, interpretation, formatting or citation behaviour. Test the copied file set as a source system rather than assuming file presence proves usable knowledge.
Create at least four reference checks:
- Positive retrieval: the relevant fact exists clearly in one approved file.
- Source conflict: two files differ and the workflow must apply a documented precedence rule or disclose the conflict.
- Absent answer: none of the files supports the requested claim.
- Scope boundary: a similarly named but unapproved or unavailable source would be needed to answer.
For each file, record title, version, date, owner, authority, sensitivity and the passages used to adjudicate the case. Where citations or source labels are required, define the exact minimum: for example, document title and section heading. Do not require a page number if the file format does not support stable pagination. Conversely, do not accept a generic “according to the documents” statement when users need to distinguish authoritative policy from commentary.
Source quality must be reviewed by a person competent to judge the material. For legal, financial, employment, security, privacy, medical, government or similarly consequential work, outputs are drafts or research aids only. A qualified human must verify current authority, applicability and consequences before any decision or action.
The failure rule is strict: if the answer cannot be supported by the approved source set, the candidate should disclose the gap, ask for an authorised source or stop according to the workflow. Fluent invention is a failure even when the invented statement appears plausible.
Test format as a contract, not an aesthetic preference
Output acceptance should distinguish structural validity, required content and optional presentation. Write a checklist that another reviewer can apply without guessing. Suitable checks include required headings, field presence, ordering, maximum scope, date convention, separation of evidence and interpretation, and omission of prohibited data.
For a sample briefing format, mandatory fields might be “Issue”, “Source position”, “Operational impact”, “Open questions” and “Human review required”. The test should use a fixed fixture and check every field. It should also include an empty-data case: if there is no supported operational impact, the plugin should use an approved “not established from supplied sources” notation rather than manufacture content merely to fill the field.
A formatting defect is mandatory or advisory according to downstream use. A missing heading may be advisory for a human-readable note but mandatory if another approved process relies on that heading. Do not infer machine compatibility from visual resemblance. Where a downstream system consumes the output, validate it separately through that system’s authorised procedure; this article does not assert compatibility with any external parser or schema.
Exercise missing tools and broken prerequisites deliberately
A positive integration test alone cannot show safe failure. OpenAI’s plugin documentation separates installation from account authorisation and notes that use can depend on plan, workspace, role, region, capability and platform. Test each prerequisite independently where the organisation can do so safely.
- Plugin installed, underlying app unavailable.
- App available, provider account not authorised.
- Provider authorised, requested source outside the user’s provider permissions.
- Read access available, write action unavailable.
- Action available but approval not granted.
- Feature available to the owner but not to an intended member role.
- Required capability unavailable on an intended platform.
For each state, ask for the dependent task and inspect both the response and the relevant authorised system record. The candidate must not claim that it read data or completed an action when the prerequisite was absent. A clear limitation message passes only if it accurately reflects the tested state and directs the user to an approved next step.
Do not place passwords, access tokens, private keys, authentication cookies or other secrets in prompts or screenshots. Use the provider’s supported authorisation path and redact identifiers from evidence. If troubleshooting requires sensitive logs, follow the organisation’s approved incident and data-handling process rather than copying them into a conversation.
Prove that unsupported actions remain unsupported
Custom actions do not transfer automatically. A skill can describe a workflow, but OpenAI’s skill documentation assigns live data, authorisation and controlled actions to an app or MCP server. If no approved execution surface replaces a former action, acceptance must test an honest refusal or manual hand-off, not pretend that prose recreates execution.
Build negative cases around the original action inventory. For example, if the old GPT submitted a record to an internal service but the candidate contains only a skill and reference files, ask it to submit a sample record. The acceptance criterion should require it not to claim submission, not to invent a confirmation identifier and not to solicit credentials in the prompt. It may prepare a draft payload or explain the approved manual procedure only if that behaviour is explicitly within scope.
If an app or custom MCP server does provide a replacement action, test dry-run, approval, cancellation, denied permission, malformed input and duplicate-request handling where those behaviours are supported and safe to test. Do not perform consequential writes against production merely to prove connectivity. Use an approved test tenant, sandbox or non-destructive fixture where available; if none exists, require a human-controlled validation method and document the residual risk.
Negative cases should confirm that a replacement denies or constrains actions the original workflow should not perform, while retaining evidence for review. An enterprise administration update on proving access and preserving audit evidence offers context for documenting Model Test access evidence controls alongside policy audit logs and group administration around the replacement.
Check direct triggers, indirect triggers and non-triggers
A migrated instruction set becoming a skill does not prove that the skill will activate in the same circumstances as the original GPT. Prepare three adjacent cases with similar vocabulary:
- A direct request naming the supported workflow.
- An indirect request describing the same goal without its formal name.
- An out-of-scope request sharing keywords but requiring a different workflow.
For example, “Create an internal policy briefing from this approved document” is a direct trigger. “Help the operations team understand what changes next month” may be an indirect trigger if the supplied material and context make the intended workflow clear. “Rewrite this marketing announcement in a friendlier tone” may contain policy-related terms but should not invoke the briefing workflow. Define these classifications before execution.
False activation and missed activation carry different costs. If activation could expose data or enable an action, prefer clarification over aggressive triggering. If the skill is read-only and low consequence, an organisation may tolerate a broader trigger provided the output accurately states its scope. Record this trade-off as an explicit owner decision rather than adjusting the criterion after seeing results.
Test intended-user installation from a clean state
The migrated personal plugin starts private because sharing settings do not carry over. Keep it private during owner testing. Publishing the source GPT for migration is not the same as public sharing, and migration does not automatically restore access for former users.
After owner acceptance, select representative intended users by role, workspace status, provider entitlement and platform. Do not use the owner account as a proxy for everyone. Each participant should begin from a documented clean state and follow the organisation’s approved installation route.
| Installation case | Procedure | Acceptance criterion | Expected | Observed | Evidence |
|---|---|---|---|---|---|
| I-01 Intended member | Share privately with one authorised member; have that person locate and install through the approved workspace route. | The named user can access only after sharing and any required workspace controls; no unrelated user is intentionally granted access. | Record role, platform, date and redacted installation evidence. | ||
| I-02 App authorisation | Have the intended user initiate the supported provider-authorisation flow without sharing credentials. | Authorisation follows the provider-supported route; installation alone is not recorded as account connection. | Record authorisation state without tokens or secrets. | ||
| I-03 Restricted member | Attempt the workflow with a role that should not use the underlying app or action. | The restriction remains effective; the plugin does not imply that installation overrides it. | Record role and limitation response. | ||
| I-04 Unshared user | Check access from an otherwise comparable account not included in the private sharing list. | The test does not establish unintended access; any discrepancy blocks wider release pending review. | Record only privacy-safe access evidence. |
Installation acceptance is not complete until the user runs one familiar task, one reference-dependent task and one applicable integration case. If platform differences matter, repeat this subset on every supported platform used by the target group. Do not advertise unsupported platforms merely because the owner’s platform works.
Require separate approval for consequential actions
Any test involving money, account changes, employment decisions, security controls, personal data, legal status, healthcare, government services or other consequential effects requires a named human reviewer and an approved non-production method wherever possible. The plugin must not be the sole decision-maker or final approver.
The test record should distinguish recommendation, draft, approval request and completed action. For example, producing a draft access-change request is not the same as changing access. The acceptance criterion should state which stage is permitted. If the approved workflow stops at a draft, any claim that the account was changed is a failure.
Where an action requires confirmation, test cancellation and denial as well as approval. Saved approval settings, where available, must not be treated as overriding provider restrictions, workspace policy or safety protections. A human reviewer should inspect the intended target, scope and consequence immediately before authorising a consequential write.
Use owner sign-off as an evidence decision
The accountable owner should not sign merely because testers report that the replacement “looks fine”. The sign-off packet should contain the frozen candidate identifier, scope, required cases, completed evidence, defects, blocked cases, exceptions, access matrix, privacy and security review status, intended-user results, rollback plan and retirement deadline relevant to the workspace notice.
Assign at least these responsibilities:
- Workflow owner: confirms that mandatory business functions and exclusions are accurate.
- Technical maintainer: confirms the candidate revision, included surfaces and defect remediation.
- Data or source owner: confirms reference authority, permitted use and retention requirements.
- Workspace administrator: confirms the intended sharing, roles and applicable workspace controls.
- Risk reviewer: reviews privacy, security, legal or other consequential implications where relevant.
- Release owner: makes the proceed, remediate, defer or retire decision.
One person may hold more than one role in a small organisation, but the record should still show which judgement that person made. For high-consequence workflows, use independent review where organisational policy requires it.
Preview privately and retain a practical rollback path
Keep the candidate private by default through configuration review, owner tests and limited intended-user preview. Expand sharing only after the previous stage passes. Public distribution, if needed, is a separate governance and publication decision; migration itself does not make the plugin public or restore the former GPT’s audience.
A preview group should be small enough to contain defects but representative enough to expose role and authorisation differences. Give participants the supported scope, known limitations, feedback route and instruction not to use the preview for consequential production decisions. Monitor defect reports by case identifier rather than relying on informal impressions.
Define rollback before preview. The FAQ says the original GPT remains usable until its applicable retirement date, but after migration it becomes read-only, cannot be deleted by its creator and later becomes inaccessible at retirement. Consequently, rollback is not an indefinite promise to reverse the migration. Before migration, preserve the published configuration, redacted test corpus, reference-file inventory, action specifications, user list and operational hand-off. During the remaining overlap period, rollback may mean withdrawing plugin sharing and temporarily directing authorised users to the still-available original GPT. After retirement, rollback must instead mean reverting to an approved manual process, another validated system or suspension of the workflow.
The release owner should choose one of four decisions:
- Proceed: all mandatory and negative gates pass, required reviews are complete and rollback is viable.
- Remediate: keep private, correct specified defects and run a new frozen candidate through affected tests.
- Defer: delay release where an applicable, approved organisational timetable permits it; a target date or requested exception is not itself evidence of an approved deferral.
- Retire the workflow: do not release a replacement when mandatory capability, access, source quality or risk controls cannot be established.
The final decision rule is conservative: uncertainty about a mandatory capability is a block, not an assumed pass. A polished answer cannot offset missing source authority, absent action controls or unintended access. Conversely, a documented difference from the old GPT is not automatically a failure if the replacement’s approved requirements are met and users are clearly told what changed.
Close the review without erasing unresolved evidence
After sign-off, retain the acceptance record according to the organisation’s approved retention policy. Preserve failed and blocked cases alongside passes; removing them would hide the basis for exceptions and future regression tests. Record the release revision, sharing scope, support owner, known limitations and the next review trigger.
Review triggers should include a skill change, replacement or addition of a reference file, app or MCP configuration change, provider-permission change, expansion to a new user role or platform, altered output contract, material workspace-policy change, or an updated OpenAI notice affecting the migration or retirement timetable. Because availability and controls can change, check the current account or workspace notice immediately before migration and release rather than relying solely on information reviewed on 1 October 2026.
For routine maintenance, rerun the smallest set that fully covers the changed component plus the familiar-task core and relevant negative cases. For example, replacing a policy reference file requires source-positive, source-conflict and absent-answer checks, as well as familiar outputs that depend on that file. Altering an action surface requires permission, denial, approval, cancellation and no-false-success cases. Changing sharing requires intended-user and unshared-user checks.
Acceptance ends only when the owner can state, with attached evidence, what the replacement supports, what it deliberately does not support, who can use it, which sources and actions require separate access, and what happens if it fails. That is a narrower claim than behavioural identity with the custom GPT, but it is the defensible basis for a controlled replacement.
Choose continuity deliberately: migration is not equivalence
The final decision is not whether a migration control completes. It is whether the replacement preserves the parts of the workflow that the organisation still needs, while exposing any loss of behaviour, access or accountability. OpenAI’s Custom GPT retirement and migration FAQ documents an intentionally uneven transfer: published instructions become a plugin skill, knowledge files are copied into plugin reference files, and connected apps are added as apps. The selected model, existing conversations, sharing settings and custom actions do not transfer automatically.
That boundary creates three materially different outcomes. A plugin may be structurally migrated but operationally unusable because its users cannot access an underlying app. It may be accessible but functionally incomplete because a custom action has no replacement. It may reproduce ordinary drafting tasks yet remain unsuitable for consequential work because source use, approval controls or negative cases have not been reviewed. Record these outcomes separately rather than assigning one overall “migration successful” label.
As of information reviewed on 1 October 2026, OpenAI describes migration and retirement milestones as planned, targeted or scheduled, with availability varying by account or workspace. Use the date analysis established earlier in this article to identify the applicable retirement date, including any qualified and approved Enterprise deferral. Do not infer that a target date means the migration experience is visible in a particular workspace. Immediately before making a change, compare the current OpenAI documentation with the notice shown to the workspace owner or administrator.
Use the following decision rule: approve replacement only when every required capability has an identified destination, every required user has verified access, and every material risk has an owner and disposition. If a requirement is merely assumed to transfer, leave the decision open. If it cannot be tested before the applicable retirement date, either remove it from scope, establish a separately approved alternative, or retire the workflow.
Record the documented risks without overstating them
A useful risk register links each risk to a documented migration boundary and a concrete verification step. It should not predict failure rates or assign numerical probabilities without organisational evidence. Suggested severity labels such as “blocking”, “conditional” and “accepted limitation” are examples of governance categories, not OpenAI product guarantees.
| Documented boundary | Resulting risk to examine | Practical evidence | Decision rule |
|---|---|---|---|
| The selected model does not carry over. | Outputs may differ because the replacement operates under the models and defaults available in the target workspace. | Record the original dependency, identify the models currently enabled for intended users, and run the approved task corpus against the replacement. | Do not approve on the basis of a model name. Approve only against task-level acceptance criteria. |
| Conversations do not move. | Examples, exceptions or informal operating knowledge may remain only in old conversations. | Identify required records and examples before retirement, then handle them under the organisation’s applicable retention and export rules. | If the workflow depends on inaccessible conversation history, pause until that dependency is removed or resolved. |
| Sharing settings do not carry over, and a migrated personal plugin starts private. | Former users may lose access, while an incorrectly rebuilt audience could receive excessive access. | Test installation and use with representative accounts for each intended role. | Neither the owner’s successful test nor former GPT access proves recipient access. |
| Custom actions do not transfer automatically. | A visible plugin may lack the live-data lookup or write operation on which the workflow depends. | Map each former action to an available app, an independently built MCP server, a manual process or retirement. | No named and validated replacement means no action parity. |
| Connected apps are added as apps. | Presence may be mistaken for provider authorisation or source access. | Verify installation, underlying-app availability, user role, provider sign-in, source permissions, supported actions and approvals separately. | Access is accepted only when all required layers work for the intended user with the minimum necessary permissions. |
| Knowledge files are copied into reference files. | File presence may be mistaken for identical retrieval, interpretation, formatting or citation behaviour. | Test questions with known answers, absent answers, conflicting passages and required source attribution. | Reject if the plugin invents support, uses an obsolete source where authority matters, or fails the defined reference requirement. |
| Instructions become a skill. | The replacement may select or follow the skill differently, especially for indirect or ambiguous requests. | Test direct triggers, indirect triggers, incomplete inputs, non-triggers and edge cases. | Approve only if the observed trigger boundary matches the documented use case and avoids unsupported actions. |
| Migration uses the latest published version. | Unpublished edits or drafts may be omitted. | Compare the intended configuration with the latest published version before migration. | Do not proceed while a required change exists only in a draft. |
For example, suppose a policy-drafting GPT used a custom action to fetch the current internal approval status for a document. The instructions and policy files may transfer, but the status lookup does not automatically become an app or MCP server. The replacement could still produce a plausible draft while silently lacking current approval data. Classify this as a blocking action gap if live status is mandatory. Classify it as a scoped limitation only if the owner removes the lookup from the supported workflow, tells users that status must be checked elsewhere, and tests that the plugin does not claim to have performed the lookup.
Some risks arise from organisational dependence rather than the migration mechanism. Examples include an undocumented reviewer rota, a file that no longer has an accountable owner, or a provider account used by one former maintainer. Keep these in the same decision record, but label them “local dependency” rather than attributing them to OpenAI. That distinction determines who can fix the issue and prevents a product migration from concealing an internal governance problem.
Turn unknowns into update gates
An unknown is not automatically a defect. It is a question whose answer could change the release decision. Manage unknowns by naming the evidence needed, the person responsible and the latest acceptable resolution point. Avoid converting “not yet checked” into either “safe” or “unavailable”.
- Confirm account-specific availability. Ask the owner to check whether migration is visible for the intended published GPT in the correct workspace. A target date in an OpenAI article does not establish account eligibility or interface availability.
- Confirm the governing notice. Have the workspace administrator record the applicable scheduled retirement date and whether an Enterprise deferral is both qualified and approved. A hoped-for or requested deferral is not an approved one.
- Confirm the published source object. Identify the latest published version and compare it with any editor draft. Resolve differences before invoking migration because unpublished edits are not the documented source for transfer.
- Confirm action replacements. For every custom action, identify whether an existing app supports the required operation or whether a separate MCP implementation is proposed. An app template is not evidence of a ready-to-use connection.
- Confirm recipient conditions. Record plan, workspace, role, region, platform, underlying-app availability and provider-account permissions where relevant. OpenAI’s plugin guidance says individual availability can depend on these conditions.
- Confirm data controls. Determine which sources can be read, which actions can write, whether approval is required, and whether synchronisation restrictions differ from direct-action restrictions.
- Confirm model conditions. Identify what is enabled in the workspace instead of assuming the original GPT’s selection persists. OpenAI’s 6 July 2026 release note describes GPT-5.5 Instant Mini as a ChatGPT fallback after GPT-5.5 Instant or Auto rate limits; it does not appear in the model picker, and that update does not affect the application programming interface or Codex.
- Confirm conversation handling. Determine whether old conversations contain records that must be retained and which account-specific archive or export process applies. The migration FAQ establishes non-transfer, not a universal export procedure.
- Confirm distribution scope. Decide whether the replacement is private, internally shared or intended for public-directory publication. Public distribution requires its own process and evidence.
A concise unknowns log might say: “Finance-status action replacement: unresolved; technical owner to determine whether an approved app supports read-only status retrieval; evidence required is provider authorisation and a representative-user access test; due before acceptance review.” This is preferable to “integration pending” because it identifies the capability, responsible role, proof and gate.
Apply a stop rule when an unknown concerns access to sensitive data or the ability to perform a consequential action. Do not resolve uncertainty by granting broader permissions, reusing another person’s account, or placing secrets in a prompt. Credentials, access tokens, personal data that is not necessary for the task, and other confidential material should remain outside prompts and test cases.
Assign one accountable migration owner without collapsing specialist review
The migration owner coordinates the decision and maintains the evidence record. That person need not personally administer every system or approve every risk. Separating accountability from specialist authority prevents two common errors: a creator approving controls they do not own, and a committee producing advice without anyone responsible for closure.
Responsibilities before migration
- Identify the custom GPT, creator, owning workspace, latest published version and applicable retirement context.
- Freeze or version the inventory used for comparison, while recording any subsequent approved change.
- Classify instructions, files, apps, custom actions, user groups and conversation-dependent knowledge.
- Nominate owners for the plugin skill, reference files, each underlying app, any MCP replacement, workspace access and user communications.
- Define the minimum supported workflow and the functions explicitly excluded from the replacement.
- Set acceptance criteria before testing so that results are not judged against convenience after the event.
- Escalate security, privacy, legal, records-management and accessibility questions to the people authorised to decide them.
Responsibilities during validation
- Ensure tests use representative but non-secret inputs, with synthetic or approved data where possible.
- Collect evidence for direct and indirect triggering, non-trigger behaviour, incomplete inputs, reference use, formatting, app access and action approval.
- Arrange clean-state tests for representative users rather than relying on the creator’s existing sessions or permissions.
- Record failed prerequisites and denied operations as evidence, not merely successful cases on the expected path.
- Keep remediation changes traceable to the test case that prompted them.
- Prevent a change to skill instructions, reference files, app configuration or access policy from bypassing required retesting.
Responsibilities at release and closure
- Obtain named sign-off from the service owner and any required specialist reviewers.
- Publish the supported-use statement, known limitations, access instructions, support route and transition date.
- Confirm who may install and use the plugin, which apps still require individual authorisation, and which actions require approval.
- Monitor unresolved issues until closure or formal acceptance, without representing accepted risk as eliminated risk.
- Retain the migration record under the organisation’s own records policy.
- Plan for the original GPT’s documented post-migration state: OpenAI says it becomes read-only after migration, cannot be deleted by its creator, and later becomes inaccessible at retirement.
A practical responsibility split might assign the workflow owner to define expected outcomes, the workspace administrator to verify installation policy and roles, the app owner to confirm provider-side permissions, and a privacy reviewer to assess the proposed data path. The migration owner then decides whether the collected evidence satisfies the release gates. They must not substitute their own judgement where a formal specialist approval is required.
Communicate the change as a capability and access change
A message that says only “the GPT is becoming a plugin” is inadequate because it suggests a rename rather than a lossy transition. Communication should tell users what remains supported, what no longer transfers, what they must do, and where uncertainty remains. It should also distinguish installation from authorisation and private availability from public publication.
Use at least four communications, scaled to the workflow’s importance:
- Owner and administrator notice. State the affected workflow, applicable workspace notice, proposed migration decision, open blockers and named owners. Send it early enough for access and action gaps to be resolved.
- Pilot-user instructions. Explain how selected users obtain access, which underlying apps require authorisation, what test tasks to perform and how to report a failure. Do not ask users to paste credentials, confidential records or unnecessary personal data into prompts.
- Release notice. State the supported tasks, excluded functions, known differences, required approvals and support route. If the former custom action is absent, say so directly.
- Retirement reminder. Remind users that old conversations do not move and that any required records must be handled through approved organisational procedures before the applicable date.
An example release notice could read:
Example, not a product-generated notice: “The Policy Briefing plugin replaces the drafting and reference functions of the Policy Briefing custom GPT. It does not carry over old conversations, the former model selection or the previous sharing list. The live approval-status action is not included; check status in the approved source system. Access to the Records app requires separate provider authorisation and remains subject to your role. Do not enter secrets or unapproved personal data. Report missing access, unsupported actions or source errors to the service owner.”
Choose the communication date by working backwards from the applicable retirement date and the time needed for remediation, not by assuming a universal product rollout date. If migration is unavailable in the workspace, communicate that status as an unresolved dependency rather than announcing a completed transition.
For an internal workflow, include an access-request route and identify who approves it. For a public-facing candidate, keep the internal migration notice separate from any directory announcement. A private migrated plugin is not publicly available merely because the source GPT was published, and “published” in the migration prerequisites does not itself mean public sharing.
Protect accessibility and user rights through explicit review
Portability includes the ability of intended users to complete the task, not merely the transfer of configuration. OpenAI’s documented migration mapping does not establish that a replacement preserves the old interaction pattern, output format or availability on every platform. Plugin capabilities can vary by plan, workspace, role, region, capability and platform. Accessibility therefore needs its own acceptance evidence.
Accessibility review procedure
- Identify essential task paths. List the shortest supported route from request to usable result, including any app authorisation or action approval step.
- Specify output requirements. Record required heading structure, plain-language constraints, table alternatives, link descriptions and any machine-readable format on which downstream users rely.
- Exercise representative interfaces. Test only the platforms and account types the organisation intends to support; do not generalise one platform’s result to all surfaces.
- Check failure communication. Ensure a denied permission, absent source or unsupported action is described clearly enough for the user to recover or seek help.
- Provide a non-plugin route where required. If a user cannot complete a necessary organisational process through the replacement, assign an accessible alternative and owner rather than treating exclusion as user error.
- Obtain human feedback. Automated format checks can identify structural omissions, but they cannot establish that a workflow is usable for every person. Include review by people who use assistive technologies and obtain their feedback before release.
For example, if users depend on a briefing in a predictable heading order, the acceptance criterion could require “Summary”, “Evidence”, “Uncertainty” and “Required decision” sections. Test the structure with a normal request, an incomplete request and a case with no supporting reference. A visually tidy response that omits uncertainty would fail the defined contract. This is a suggested review method, not a claim about how any particular plugin will respond.
Rights and access safeguards should cover both exclusion and overreach. Rebuild access from named organisational roles and legitimate task needs; do not copy an old audience from memory. Verify that an intended user can install or invoke the plugin, access the necessary underlying app and use only the required sources and actions. Equally, verify that a user outside the intended role cannot reach restricted functions through the plugin.
Provider-side rights remain decisive. An app-permission setting cannot connect an account, grant provider access, change source-system rights or override workspace policy. Saved action approval does not override restrictions or safety protections. Consequently, a successful owner test does not prove that another employee can read the same records, and workspace permission does not itself grant rights inside the source system.
Any workflow involving security, privacy, money, employment, government services, healthcare, legal status or similarly consequential decisions requires qualified human review. The plugin may assist with retrieval, drafting or structured analysis where permitted, but its output must not be treated as the sole authority for a decision affecting a person’s rights, access, finances, employment or safety. Define who reviews the source evidence, who may approve an action and how a person can challenge or correct an error.
Use a data-minimisation rule for validation: include only the information necessary to test the requirement. Keep passwords, authentication tokens, private keys and other secrets out of prompts. Avoid copying production personal data into test cases when synthetic, redacted or separately approved data can exercise the same behaviour. If realistic sensitive data is genuinely necessary, pause until the authorised privacy and security reviewers approve the method and environment.
Keep public-directory publication outside the migration acceptance decision
Private replacement readiness and public distribution answer different questions. The first asks whether defined users can perform an approved workflow under the intended controls. The second asks whether a packaged plugin should enter a public directory through OpenAI’s separate submission, checks, review, approval and publication process. Passing an internal acceptance review neither submits nor publishes a plugin.
Run a separate public-release review only after the private capability is stable:
- Confirm that public distribution is genuinely required rather than assuming broader reach is beneficial.
- Define what information, skills, reference material and app capabilities may be exposed to an unknown audience.
- Remove internal-only instructions, confidential reference files, organisation-specific endpoints and secrets from the public package.
- Review the compressed-archive submission and associated configuration against the current OpenAI developer requirements.
- Complete the applicable checks and review process, then make a separate decision to publish after approval.
- Retest the approved package and public description if the submission process requires material changes.
- Maintain an owner and support route for updates, access problems and withdrawal decisions.
Do not describe directory review, approval or verification as a security certification. It does not replace the publisher’s security, privacy, legal, accessibility or vendor review, nor does it prove that an underlying provider grants a user access. A directory listing also does not make every capability available on every plan, platform, region or workspace.
For example, an internal plugin may include a skill for drafting procurement notes and an app that reads a restricted contract repository. A public version cannot be assumed to inherit that app access safely or lawfully. The publisher would need to decide whether the public package should omit the integration, use a separately designed public data source, or not be published. “The internal plugin works” is irrelevant evidence for that distribution decision.
Apply a four-outcome decision rubric
Use a named outcome rather than a single numerical score. Scores tend to conceal blockers by allowing several minor strengths to offset one unacceptable access or action gap. The following rubric treats mandatory requirements as gates.
Proceed
Choose Proceed when the correct published GPT is eligible and the migration path is available; all required capabilities have documented destinations; representative users have verified access; action and reference tests meet the predefined criteria; sharing has been rebuilt deliberately; communications and support are ready; and all required specialist reviewers have signed off.
Example: instructions and approved files cover the whole supported workflow, no custom action is required, intended users can access the private plugin, and the reference tests satisfy the organisation’s source and format rules. The record may still contain minor limitations, but none contradicts the supported-use statement.
Proceed with an explicit limitation
Choose Proceed with an explicit limitation when a non-essential capability is absent or behaves differently, the supported scope can be narrowed safely, users are told what changed, and tests show that the plugin does not claim the excluded capability.
Example: a convenience action that formatted a draft in an external system has no replacement, but drafting and reference use remain valid. The owner removes “publish to system” from the supported workflow, provides a documented manual hand-off, and tests that the plugin does not say publication occurred. Do not use this outcome if the missing function is necessary for safety, legal compliance, approval or record integrity.
Remediate before release
Choose Remediate before release when the intended replacement is feasible but evidence is incomplete or a required control fails. Typical triggers include an unpublished source change, missing provider authorisation, an inaccessible reference file, over-broad action rights, failed non-trigger behaviour, or an unresolved accessibility barrier.
Assign each issue an owner, evidence requirement and retest scope. If an app permission changes, rerun affected access, source and action cases; do not rerun only the original failed prompt. If skill instructions change, repeat direct, indirect, incomplete-input and non-trigger cases because the trigger boundary may have shifted.
Retire or replace the workflow by another route
Choose Retire or replace by another route when a mandatory custom action has no acceptable replacement, the required data access cannot be granted appropriately, the workflow depends on conversation history that cannot be reconstructed, or the residual risk cannot be approved. This outcome is not a failed migration if it prevents an unsupported process from being presented as continuous.
A separate technical project may later produce an MCP-backed replacement, but do not approve the plugin on the expectation that one will exist. Until implementation, provider authorisation, permissions and action behaviour are validated, record the capability as unavailable.
Decision questions in order
- Is the source definite? If the owner, workspace or latest published version is uncertain, stop and identify it.
- Is migration actually available? If not, retain the issue as an account-specific dependency and follow the applicable notice rather than inferring availability from a target date.
- Does every mandatory function have a destination? Instructions, files and apps may have documented mappings; selected models, conversations, sharing and custom actions require different treatment.
- Can intended users obtain appropriate access? Test installation, app availability, role, provider authorisation, source rights and action controls separately.
- Does the candidate meet task and negative-case criteria? Familiar prompts alone are insufficient; include harder, incomplete, non-trigger and unsupported-action cases.
- Are consequential effects human-controlled? Require authorised human review and approval for high-impact decisions and actions.
- Can known differences be communicated truthfully? If a limitation would make the supported-use statement misleading, do not release.
- Is private readiness being confused with public publication? If public distribution is required, open the separate directory review.
The final decision record should name the outcome, date, accountable owner, supported scope, excluded capabilities, intended audience, evidence set, unresolved items, specialist approvals and next review trigger. It should also state that completion of migration is not evidence of identical behaviour or preserved access.
A compact example is: “Decision: remediate before release. Supported candidate scope: drafting from approved reference files. Blockers: provider authorisation not verified for regional reviewers; former submission action has no approved replacement. Required evidence: clean-state role tests and a decision to remove or rebuild submission. Public-directory review: out of scope. Decision owner: workflow service owner.” This example records no invented product result; it shows how to make the judgement auditable.
Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!
Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.
Useful Links
- OpenAI: Custom GPT retirement and migration FAQ
- OpenAI: Plugin documentation
- OpenAI Developers: Build plugin skills
- OpenAI: Managed-workspace administration guidance
- OpenAI: Managing app permissions in ChatGPT
- OpenAI: ChatGPT release notes
- OpenAI: Codex models
