How to Run Deep Research in ChatGPT Work and Codex: Source Controls, Live Steering, Citations, and Editable Deliverables

How to Run Deep Research in ChatGPT Work and Codex: Source Controls, Live Steering, Citations, and Editable Deliverables
How to Run Deep Research in ChatGPT Work and Codex: Source Controls, Live Steering, Citations, and Editable Deliverables

What changed on September 9: Deep Research moved beyond ordinary Chat into Work and Codex

OpenAI’s September 9 release notes state that Deep Research is now available in ChatGPT Work and Codex, expanding the feature from a research report experience in ordinary Chat into longer-running work surfaces where users may be assembling deliverables or working with software repositories. The practical change is not that Deep Research became an unlimited autonomous analyst; the change is that eligible users can now ask for deep research inside the same Work or Codex context where they are already managing a multi-step task, steering scope while the task runs, and requesting an editable artifact when the required tools and permissions are available.

OpenAI describes Deep Research as a workflow that can research across the web, user files, and supported connected apps, then return a structured report with citations or source links. In Work and Codex, OpenAI says users can start with @Deep Research in Work or explicitly ask for deep research, add instructions while the research is running, and request outputs such as a document, presentation, spreadsheet, or Site when the needed tools are available. Those boundaries matter: the feature can only use sources and output paths that are available and authorized in the current task, and a generated file or link still has to be opened, inspected, and verified by the user.

This tutorial treats the September 9 expansion as an operating procedure for evidence-safe work. The goal is to help a founder produce an investor diligence memo without mixing public claims and private assumptions, help an enterprise administrator constrain research to approved knowledge sources, help a research lead request a literature-style evidence table with citations, help a security team prevent unsupported claims from entering a risk register, and help a Codex user delegate repository-adjacent research without granting access beyond the current environment’s permissions.

When to use Deep Research instead of search or standard chat

Use ordinary search when the task is a narrow lookup: a product page, a release date, a statutory text, a vendor support article, or a current fact that you can verify directly. A search-first workflow is usually faster and easier to audit when the decision depends on one or two primary sources, because you can inspect those sources yourself and avoid overbuilding a report around a simple retrieval task.

Use standard chat when the model can work from context you already provide, such as rewriting a policy, comparing two uploaded excerpts, drafting an email from notes, brainstorming interview questions, or explaining a technical concept. Standard chat is the better fit when you need reasoning, formatting, or synthesis over known material rather than a broad evidence-gathering pass across the web, uploaded files, and authorized connected sources.

Use Deep Research when the task has multiple evidence paths, a decision audience, a defined time horizon, and a deliverable that must carry citations or source links. Strong candidates include market maps, vendor scorecards, due-diligence checklists, customer-research syntheses, policy briefs, rollout plans, risk registers, technical trade-off memos, and source-backed executive summaries. OpenAI Academy’s Deep Research guidance recommends specifying the goal, scope, timeframe, output format, permitted sources, and evidence boundaries up front, then reviewing the live plan and validating key claims before reuse.

Do not use Deep Research as a substitute for professional review or source verification. OpenAI’s Deep Research documentation says citations improve traceability but do not guarantee correctness, and it instructs users to verify underlying sources, especially for decisions, external publishing, or high-stakes topics. In practice, that means a citation can show where a claim came from, but it does not prove the cited page is current, complete, unbiased, legally sufficient, or correctly interpreted.

The core contract: outcome, context, sources, constraints, and output

The safest way to start Deep Research in Work or Codex is to treat your prompt as a contract with five parts: outcome, context, sources, constraints, and output. This contract reduces ambiguity before the system asks clarifying questions, makes live steering easier, and gives you a checklist for verifying the final artifact. If one part is missing, Deep Research may still proceed, but it has more room to infer priorities, include irrelevant evidence, or produce a report that looks polished while failing the actual business or technical decision.

Contract element What to specify Operational example Verification question
Outcome The decision, deliverable, or action the research must support. “Prepare a source-backed vendor shortlist for a 90-day data catalog pilot.” Does the final answer help someone decide, approve, reject, or plan?
Context Audience, business constraints, technical environment, geography, timeframe, and known assumptions. “Audience is the CIO and security lead; company uses Microsoft identity tooling; focus on North America; ignore vendors without enterprise SSO documentation.” Did the report use the context as a filter rather than as decorative background?
Sources Allowed websites, uploaded files, authorized connected apps, preferred primary sources, and excluded sources. “Use vendor documentation, public pricing pages where available, uploaded procurement notes, and security whitepapers; do not rely on unsourced blog summaries.” Can each material claim be traced to an allowed source or clearly labeled as inference?
Constraints Evidence boundaries, exclusions, confidentiality rules, citation rules, uncertainty handling, and required human-review points. “Separate evidence from inference; flag unsupported claims; do not include confidential customer names in any externally shareable version.” Did the output preserve uncertainty instead of converting weak evidence into confident recommendations?
Output Format, structure, length, artifact type, tables, appendix, spreadsheet columns, slide outline, or editable file request. “Return an editable memo plus a comparison table with criteria, evidence, source links, risk notes, and follow-up questions.” Does the generated file or link exist, open correctly, and contain the requested structure?

The contract should be explicit about evidence hierarchy. For example, a market-sizing task might prefer regulator filings, company reports, and primary datasets over news articles and analyst commentary; a security evaluation might prefer vendor documentation, security advisories, and policy text over forum posts; a repository research task in Codex might prefer files already available in the workspace and official library documentation over broad web commentary. The point is not to ban every secondary source, but to make the model distinguish evidence, inference, uncertainty, and unsupported claims.

Starter prompt: Work deliverable with source boundaries

Use @Deep Research to prepare a decision-ready brief.

Outcome: Help the leadership team decide whether to pursue [decision].
Audience: [roles and level of technical depth].
Timeframe: Focus on evidence from [time period], unless older primary sources are still authoritative.
Context: [business, product, jurisdiction, technical environment, known assumptions].
Sources: Use [allowed uploaded files], [authorized connected sources if available], and public web sources from [preferred source types or domains]. Exclude [source types or domains].
Constraints: Separate evidence from inference. Flag unsupported claims. Preserve citations or source links for material claims. Identify items requiring human, legal, security, financial, or domain-expert review.
Output: Produce an editable [document/presentation/spreadsheet/Site if available] with an executive summary, findings, evidence table, risks, open questions, and recommended next steps.

This prompt is intentionally more rigid than a casual chat request because Deep Research may run across multiple retrieval and synthesis steps. The structure gives you natural control points: approve or revise the research plan, answer clarifying questions, interrupt if it is following a weak evidence path, add a missing source, narrow scope if it is becoming too broad, and request a revised artifact if the first output is not aligned with the decision.

How Chat, Work, and Codex differ for Deep Research

OpenAI’s documentation distinguishes Chat, Work, and Codex as separate experiences. Chat is for quick conversation and general interaction; Work is designed for longer multi-step work and finished deliverables; Codex is for software development and repository work. That distinction affects where you start, what context is available, what allowance is consumed, and what kind of final artifact you should expect.

Surface Best use of Deep Research Allowance or credit behavior described by OpenAI Important access boundary
Chat General research reports, plan review, public-web and uploaded-file research, and supported connected-app research when available. OpenAI says Deep Research in Chat uses a plan-dependent task allowance, and the Chat Deep Research limits remain unchanged by the Work and Codex expansion. In Chat, connected apps are used through supported read actions for research rather than arbitrary write actions, and access still depends on plan, workspace, role, session tools, provider permissions, and connected-account permissions.
Work Longer business, research, operational, and knowledge-work deliverables that may become editable documents, presentations, spreadsheets, or Sites when tools are available. OpenAI says Deep Research in Work consumes the existing Work/Codex allowance or credits and does not consume the separate Chat Deep Research task allowance. Work availability is documented for Plus, Pro, Business, Enterprise, and Edu users with Work access on web, desktop, iOS, and Android, subject to the applicable plan, workspace, region, and settings.
Codex Software-development and repository-adjacent research, such as investigating APIs, summarizing implementation options, producing technical comparison notes, or preparing engineering decision records from available context. OpenAI’s Work and Codex guide says Work and Codex share the relevant allowance or credit pool; task cost can vary by task, input and output size, reasoning settings, and mode. OpenAI states Codex is not a selectable experience on web or mobile, although supported desktop Codex chats can be accessed through the Remote tab in the mobile app.

The allowance distinction is important for teams that budget Deep Research separately from interactive chat. If a user runs Deep Research in Chat, it uses the Chat Deep Research task allowance described for that plan. If the user runs Deep Research in Work, OpenAI says it uses the existing Work/Codex allowance or credits instead. That means a research-heavy Work session can affect the same pool used by other Work or Codex tasks, while leaving the separate Chat Deep Research allowance unchanged.

The surface distinction is also important for administrators. OpenAI’s Deep Research documentation says starting Deep Research does not grant access to additional files or apps, and that workspace, role, session-tool, provider, and connected-account permissions continue to apply. For Enterprise and Edu, OpenAI notes that web search must be enabled for Deep Research access and that administrators govern access through role-based controls. A user prompt cannot override those permissions, and a report should not be treated as complete if the necessary source systems were unavailable or unauthorized.

The first operational rule: source access is not source authority

Deep Research can only work with sources it can access, but access does not make a source authoritative. A copied slide deck in a workspace may be outdated, a vendor page may omit limitations, a public article may summarize a primary document incorrectly, and an internal note may reflect a single team’s assumption rather than company policy. Your prompt should therefore specify both allowed sources and preferred evidence types, then require the final report to label conflicts, gaps, and uncertainty.

A practical source-control rule is to separate “may use,” “should prioritize,” and “must not use.” “May use” defines the accessible universe, such as uploaded files and authorized connected apps. “Should prioritize” defines the evidence hierarchy, such as primary documentation, official filings, standards text, or repository files. “Must not use” defines hard exclusions, such as confidential customer names in an external draft, outdated internal planning notes, or commentary sources that cannot support a high-stakes recommendation.

Operational warning: Citations are traceability aids, not correctness guarantees. Before a Deep Research output is used for procurement, policy, security, legal, financial, medical, academic, or public-facing work, assign a human reviewer to open the cited sources, confirm that the cited material supports the claim, check whether newer evidence changes the conclusion, and verify that sensitive or restricted material is not included in the artifact.

The opening move, then, is not “research this topic.” The opening move is a controlled research contract: define the decision, identify the audience, declare the evidence boundary, constrain the sources, preserve uncertainty, and request an output that can be inspected. In the remaining sections of this tutorial, that contract becomes a step-by-step workflow for starting Deep Research in Work and Codex, reviewing the plan, steering the run, requesting editable deliverables, and validating the final citations and files before anyone relies on them.

Build the source map before you start: files, apps, web boundaries, exclusions, and review gates

How to Run Deep Research in ChatGPT Work and Codex: Source Controls, Live Steering, Citations, and Editable Deliverables — first editorial explainer visual

Deep Research becomes more reliable when you prepare its source environment before you ask for a report. OpenAI’s Deep Research guidance describes a flow in which the user defines the outcome and report structure, provides context and sources, answers clarifying questions, monitors progress, steers the task while it runs, and receives a structured report with citations or source links. For Work and Codex users, that means the setup step is not administrative overhead; it is the control surface that tells the system which evidence is in scope, which evidence is out of scope, and which permissions already exist for the current task.

The safest starting point is a source map: a compact inventory of the materials Deep Research may use, the materials it should prefer, and the materials it must avoid. This is especially important in Work and Codex because the task may combine public web research, uploaded files, and supported connected apps, while still remaining bounded by workspace controls, role-based access control, session tools, provider permissions, and the user’s connected account. Starting Deep Research does not grant new access to files or apps, and it does not turn a disconnected provider into an authorized source.

Step 1: Create a source map that separates evidence, context, and forbidden material

A useful source map has four columns: source category, exact materials, intended use, and handling rule. The goal is to prevent Deep Research from treating every attached document or searchable page as equal. For example, a board-approved strategy deck might be allowed as background context but not as citation evidence; a signed customer interview transcript might support findings but require anonymization; an old competitive matrix might be useful only to identify questions that need fresh verification.

Source category What to list before launching How Deep Research should treat it Operational warning
Uploaded files File names, versions, dates, owners, and whether each file is final, draft, or archival Use as primary evidence, background context, or a source of questions to verify elsewhere Do not assume a file is current because it is attached; state the freshness rule explicitly.
Supported connected apps Which app accounts are connected, what workspace or folder is relevant, and what records are in scope Use only through supported actions and only where the current user and provider permissions allow access Connecting an app does not override provider-side permissions, workspace RBAC, or account-level limits.
Public web Preferred official sites, industry sources, documentation pages, public filings, or other open sources Use for current facts, public claims, market context, and citation-backed verification Web search must be enabled for Enterprise and Edu Deep Research access, according to OpenAI.
Excluded sources Competitor blogs, stale drafts, unsupported directories, unaudited exports, or sources outside the decision scope Do not use, or mention only as excluded context if needed for transparency An exclusion should be specific enough to enforce; “avoid unreliable sources” is too vague.

Recommendation: Treat “permitted” and “preferred” as different instructions. A permitted source may be used if it is relevant; a preferred source should be checked first or weighted more heavily for particular claims. A public product page, a signed contract, and an internal roadmap may all be permitted, but only one may be authoritative for pricing, commitments, or planned features. Tell Deep Research which source wins when documents conflict.

Source map template

Goal:
- Produce a decision-ready research deliverable for [audience] about [decision or question].

Primary sources:
- Uploaded file: [file name], [date], [owner], use for [purpose].
- Connected app source: [app/workspace/folder or record type], use for [purpose] if available through current permissions.
- Public web source class: [official documentation, public filings, regulator pages, standards bodies, vendor pages, peer-reviewed literature, or other class].

Preferred evidence hierarchy:
1. [Highest-authority source type]
2. [Second source type]
3. [Background-only source type]

Excluded sources:
- Do not use [source class] because [reason].
- Do not use documents older than [date] unless explicitly labeled historical.

Conflict rule:
- If sources disagree, quote or summarize both, cite both, and mark the conflict for human review rather than choosing silently.

Step 2: Upload files deliberately, not in bulk

File uploads work best when each file has a defined role in the research plan. OpenAI says Deep Research can use uploaded files by default in Chat, and Work/Codex tasks can use sources available in the current task subject to the product’s controls and permissions. The practical rule is to attach only the materials that the researcher would reasonably hand to a human analyst, then describe how each file should be used. Bulk uploading a folder of mixed drafts, exports, screenshots, and stale notes increases the risk that weak material will be cited or that the final report will blur current and historical facts.

Before uploading, rename or annotate files so the task can distinguish version and status. A file labeled “customer_interviews_final_2026_Q3” is easier to handle than a file labeled “notes.” If the filename cannot be changed, add a source note in the prompt that states the file’s date, owner, intended use, and any limitations. This is not cosmetic; it gives Deep Research a basis for deciding whether to cite the file, treat it as context, or flag it as uncertain.

  1. Attach only in-scope files. Exclude drafts that should not influence the report, privileged materials that should not be summarized, and exports whose provenance cannot be checked.
  2. State the authority level. Mark each file as “primary evidence,” “background context,” “historical reference,” or “question source.”
  3. State sensitivity handling. If a file contains confidential, personal, contractual, or unreleased information, instruct the system not to quote it in externally shareable text without a separate review.
  4. State citation expectations. Ask for file-backed claims to cite the file name or available source link, and ask for uncertain claims to be labeled rather than smoothed over.

Example instruction: “Use the uploaded customer-interview synthesis as qualitative evidence for themes, but do not treat it as a statistically representative sample. Use the uploaded pricing memo only as internal context; any external pricing statement must be verified against a public or contractually approved source before inclusion.” This kind of instruction prevents a polished report from turning internal working material into unsupported external claims.

Step 3: Authorize supported apps before the research run, then verify the permission boundary

Deep Research can use supported connected apps when available, but OpenAI states that starting Deep Research does not grant access to additional files or apps. The connected provider, the current user’s account, workspace policy, role permissions, and session tool availability continue to govern what can be read or used. In Chat, OpenAI describes connected-app research as using read actions rather than arbitrary app write actions. In Work or Codex, output to a connected app depends on supported capabilities and permissions; users must not assume that a requested document, presentation, spreadsheet, or Site can be created in every connected destination.

For administrators and security teams, the key distinction is between authorization and delegation. Authorizing a supported app makes certain app data available within the limits of the provider and workspace; it does not delegate the user’s entire identity for unrestricted operations. Deep Research should be treated as a research task operating within existing controls, not as a permission escalator. If a user cannot open a restricted folder, repository, ticket queue, or record in the provider itself, starting a research run should not be expected to make that content available.

Permission layer What it controls What to check before launch
Workspace RBAC Whether the member’s role can use Deep Research, Work, Codex, web search, browser/network access, connected apps, or particular models/tools Confirm the member’s role and workspace policy allow the intended research surface and source type.
Provider permissions Which documents, records, folders, repositories, tickets, or app objects the connected account can access Open the provider directly as the same user and confirm the target material is accessible there.
Session and task tools Which tools are enabled for the current Work or Codex task, including web, files, browser/network access, or app actions Check that the active task has the tools needed for the stated source map.
Output capability Whether a requested document, presentation, spreadsheet, Site, or app artifact can be produced After completion, open the generated file or link and confirm it exists, loads, and contains the expected content.

OpenAI’s Work and Codex guidance also distinguishes Work from Codex: Work is for longer multi-step work and finished deliverables, while Codex is for software development and repository work. That distinction matters during source setup. A research request about an implementation plan may belong in Work if the output is an executive document; a request that inspects a repository, summarizes architectural decisions, or prepares a code-adjacent migration checklist may belong in Codex, subject to the organization’s Codex Local, browser, network, and repository controls.

Step 4: Define web requirements, domain priorities, and exclusions where the interface supports them

OpenAI states that, in Chat, users can review and modify a proposed Deep Research plan and restrict research to specified websites or prioritize selected sites while allowing full-web search. For Work and Codex, OpenAI describes users adding instructions to steer or revise scope. Because controls can vary by surface, account, workspace, and enabled tools, write your domain policy in the prompt even when an interface-level selector is available. The written policy becomes part of the research contract and gives you a clear basis for interrupting the run if the plan drifts.

Use a three-tier web policy. Tier one is required sources: the task must check them if the claim depends on public evidence. Tier two is preferred sources: the task should use them first but may search wider if they are incomplete. Tier three is excluded sources: the task should not use them for evidence, or should use them only to identify claims that require verification elsewhere. This structure is more precise than saying “use authoritative sources,” because authority depends on the question being answered.

Web source policy template

Required public sources:
- For product capability claims, use official product documentation or release notes when available.
- For legal, regulatory, or compliance claims, use primary regulator, statute, official guidance, or counsel-approved material.
- For market claims, distinguish vendor claims, analyst interpretation, customer evidence, and your own inference.

Preferred public sources:
- Prefer sources published within [timeframe] unless the topic is historical.
- Prefer primary sources over summaries.
- Prefer sources that identify date, authoring organization, and evidence basis.

Excluded public sources:
- Do not cite anonymous reposts, unverified social posts, copied slide decks, or pages that cannot be traced to an accountable publisher.
- Do not use paywalled snippets as evidence unless the full source is available for review.
- Do not rely on marketing comparisons unless they are labeled as vendor claims.

Enterprise and Edu users should also account for OpenAI’s web-search requirement for Deep Research access. If web search is disabled by policy, a task that depends on current public evidence may be unable to complete as requested. The correct response is not to ask the model to “do its best” from memory; the correct response is to change the research design, enable the appropriate tool through admin-approved channels, or limit the deliverable to uploaded and connected sources with an explicit evidence limitation.

Step 5: Add exclusions that protect confidentiality and reduce citation noise

Exclusions should cover more than domains. A mature research brief excludes stale timeframes, non-authoritative file versions, sensitive fields, unsupported jurisdictions, speculative claims, and output formats that create a false impression of certainty. If the task concerns a procurement decision, you may exclude vendor-authored benchmark claims unless independently verified. If the task concerns internal strategy, you may exclude unpublished customer names from the final artifact while allowing anonymized themes. If the task concerns code, you may exclude secrets, credentials, production data, or unrelated repositories from the research scope.

Recommended exclusion language: “Do not include confidential customer names, personal data, credentials, contract terms, or unreleased roadmap details in the final deliverable. If any such material is necessary to explain a conclusion, replace it with a neutral descriptor and add a private review note identifying what must be checked by the owner before sharing.” This instruction supports the OpenAI Academy recommendation to recheck citations, remove sensitive information, confirm rights and permissions, avoid long copied passages, and retain citations before external sharing.

Decision rule: If a source cannot be quoted safely, cited accurately, and reviewed by an accountable owner, keep it out of the externally shareable artifact. Use it only as internal context or as a prompt for further verification.

Step 6: Answer clarifying questions as control checkpoints, not interruptions

OpenAI’s Deep Research guidance includes clarifying questions as part of the normal workflow. Treat those questions as a quality gate. If the system asks about audience, timeframe, source preference, output format, or decision criteria, answer with enforceable rules rather than broad preferences. A weak answer such as “make it concise” gives little operational guidance; a stronger answer says “write for a CFO and security lead, limit the executive summary to five bullets, include a risk register, and separate vendor claims from independently verifiable evidence.”

Clarifying questions are also the right time to tighten permissions. If Deep Research proposes using the open web but the decision must be based only on uploaded diligence files, say so before the run begins. If it proposes using connected app records that may include sensitive personal information, narrow the folder, project, ticket type, or date range. If it asks whether to create a presentation, confirm whether the tool and destination are supported in the current Work/Codex context, and require final verification that the generated artifact opens successfully.

Clarifying-answer pattern

Audience:
- [Who will read the deliverable and what decision they will make]

Timeframe:
- Use evidence from [date range]. Treat older material as historical unless it is a stable policy or long-lived contract.

Evidence boundary:
- Use uploaded files [A, B, C], connected source [X] if available, and public web sources that meet the web policy.
- Do not use [excluded materials].

Output:
- Produce [document/presentation/spreadsheet/Site if supported] plus a claim-to-source appendix.
- Mark unsupported or uncertain claims instead of filling gaps.

Review:
- Before finalizing, list assumptions, unresolved conflicts, missing evidence, and items requiring human or expert review.

Step 7: Review the research plan before execution and revise it while the run is live

Before Deep Research proceeds, inspect whether the proposed plan matches the source map. The plan should identify the question, subquestions, source classes, evidence hierarchy, output structure, and verification steps. If it omits a required source, adds an excluded source, or treats background material as primary evidence, revise the plan before the system invests time in the wrong path. In Chat, OpenAI documents plan review and modification, including site restrictions or site prioritization where available; in Work and Codex, use instructions to steer or revise the scope.

Live steering is not a last resort. OpenAI says users can monitor progress and interrupt or steer Deep Research while it runs. Use that capability when the visible trajectory shows drift: the task is spending time on generic market background instead of your uploaded evidence, it is expanding to jurisdictions you excluded, it is treating vendor claims as neutral facts, or it is building the wrong artifact type. A timely correction is cheaper than repairing a polished but mis-scoped report at the end.

Plan symptom Correction to send during plan review or live steering
The plan relies on broad web search when internal evidence is required “Use the uploaded files as the primary evidence base. Use web search only to verify public facts and clearly label any web-derived claim.”
The plan ignores a connected source “If the connected source is available under my current permissions, include it in the evidence review; if it is unavailable, state that limitation.”
The plan treats all citations as equivalent “Rank citations by the evidence hierarchy and mark vendor claims, internal assumptions, and independent verification separately.”
The plan proposes an unsupported or uncertain output destination “Prepare the deliverable in a supported editable format if available, and include a final check that the file or link exists and opens.”

The final setup checkpoint is allowance and metering. OpenAI states that Deep Research in Work uses the existing Work/Codex allowance or credits and does not consume the separate Chat Deep Research task allowance. A complex source map with many files, broad web research, connected sources, and a polished deliverable may be more expensive in task time or credits than a narrow brief. Narrowing the question before launch is therefore not only a quality control; it is a usage-control practice for teams that need predictable research operations.

Once the source map, permissions, exclusions, clarifying answers, and plan are aligned, start the run with a final instruction that preserves uncertainty. Ask Deep Research to cite sources, but do not treat citations as proof. OpenAI explicitly advises users to verify underlying sources, especially for decisions, external publishing, or high-stakes topics, and notes that citations improve traceability but do not guarantee correctness. Your setup should therefore require a final appendix that separates evidence-backed findings, inferences, unresolved conflicts, missing evidence, and items requiring human review.

Steer the live run before it turns weak evidence into a polished artifact

How to Run Deep Research in ChatGPT Work and Codex: Source Controls, Live Steering, Citations, and Editable Deliverables — second editorial workflow visual

OpenAI’s Deep Research guidance describes monitoring progress and interrupting or steering a task as part of the normal workflow, not as an emergency-only action. Treat the live run as a research review meeting: inspect what it appears to be searching, watch for early framing errors, and correct scope before the system spends Work/Codex allowance or credits assembling an answer around the wrong premise.

In Work and Codex, steering is especially important because the output may become an editable document, presentation, spreadsheet, or Site when the relevant tools and permissions are available. A well-formatted deliverable can still be wrong if the evidence boundary was too broad, if the citations do not support the claims, or if the task used accessible but low-quality sources because you did not prioritize better ones.

Use a live progress review checklist instead of waiting for the final report

During the first visible phase of the run, evaluate whether the system is following the outcome contract you gave it: audience, timeframe, geography, allowed sources, exclusions, and final artifact type. If the visible plan or progress summary drifts from those constraints, interrupt with a correction rather than hoping the final answer will self-correct.

What to inspect during the run Failure pattern to catch early Steering instruction to send
Search targets and source classes The run relies on general web commentary when you asked for primary filings, product documentation, or internal files. “Pause broad commentary. Prioritize primary sources first, then use analyst or media sources only for interpretation clearly labeled as secondary.”
Timeframe The run includes stale material outside the decision window or misses recent source updates. “Restrict the evidence review to materials published or updated within the specified timeframe, and list any older sources only as background.”
Geography or market segment The research drifts into global or adjacent-market claims when the decision is regional or segment-specific. “Narrow the analysis to the named geography and segment. Move other markets to an appendix labeled ‘out of scope’ only if they explain a material contrast.”
Connected and uploaded sources The run ignores uploaded context or assumes access to apps that were not actually authorized. “Use only files and connected sources visible in this task. If a needed source is unavailable, record it as an evidence gap rather than inferring from memory.”
Artifact structure The run is producing a narrative essay when the team needs a spreadsheet, scorecard, or slide-ready executive brief. “Keep researching, but restructure the final output as the requested artifact with the acceptance checks already specified.”

Do not use steering to smuggle in new access assumptions. OpenAI’s documentation states that starting Deep Research does not grant additional file or app access; workspace, role, connected-account, provider, and session-tool permissions still apply. If the run cannot reach a needed source, the correct instruction is to record the missing source and explain the impact on confidence, not to imply that Deep Research can bypass the permission boundary.

Interruption prompts for scope correction

A useful interruption is short, explicit, and operational. State what is wrong, what should change, and how the final output should record the change. Avoid vague commands such as “be more accurate” or “focus better,” because they do not tell the research agent which sources, claims, or sections to rework.

Live steering prompt: narrow an overbroad run

Pause and correct scope before continuing.

The current research appears to be treating this as a general market overview. The decision we need is narrower: [decision], for [audience], covering [geography/segment], during [timeframe].

From this point:
1. Exclude sources outside [geography/segment] unless they explain a direct dependency.
2. Separate primary evidence from secondary commentary.
3. Mark any already-collected out-of-scope evidence as background, not as support for the recommendation.
4. In the final deliverable, include a short "Scope corrections made during research" note.

Use a different interruption when the model is chasing a tempting but irrelevant thread. For example, a founder evaluating customer onboarding friction may not need a full survey of enterprise procurement unless procurement is part of the onboarding path. The correction should preserve useful discoveries while stopping unnecessary expansion.

Live steering prompt: stop an irrelevant branch

Stop expanding the [topic/branch] thread unless it directly affects [decision criterion].

Keep only facts from that branch that change one of these outputs:
- the recommended decision,
- a risk rating,
- an implementation dependency,
- a citation-backed assumption,
- or a question requiring expert review.

Move all other material to "researched but not used" or omit it from the final artifact.

Ask for contradiction hunting before synthesis

Contradiction hunting should happen before the final draft, because contradictory evidence is easiest to hide once the report becomes a smooth narrative. Ask Deep Research to deliberately search for sources that challenge the emerging conclusion, separate true contradictions from definitional differences, and identify which claims should be downgraded or qualified.

Live steering prompt: contradiction hunt

Before writing the final recommendation, perform a contradiction pass.

For each major claim you plan to make:
1. Find evidence that supports it.
2. Find evidence that disputes, limits, or narrows it.
3. Distinguish direct contradiction from differences in timeframe, geography, customer segment, measurement method, or terminology.
4. If the contradiction is unresolved, downgrade the claim and label it as uncertain.
5. Add a "Contradictions and unresolved conflicts" section with citations or source links.

This instruction is particularly valuable for competitive analysis, policy interpretation, vendor selection, due diligence, and research summaries based on heterogeneous sources. Citations improve traceability, but OpenAI explicitly warns that citations do not guarantee correctness; contradiction hunting helps you test whether the cited evidence actually supports the conclusion.

Request an evidence-gap report while there is still time to fix it

An evidence-gap request asks the system to stop optimizing for completeness of prose and start auditing completeness of proof. The goal is to identify unavailable sources, missing perspectives, weak citations, and claims that require human review before the final artifact is used in a decision, customer meeting, board packet, or external publication.

Live steering prompt: evidence-gap audit

Run an evidence-gap audit before final drafting.

Create a table with these columns:
- Claim or section
- Evidence currently available
- Missing evidence or unavailable source
- Why the gap matters
- Confidence impact: high, medium, or low
- Recommended human follow-up

Do not fill gaps with assumptions. If a connected app, file, or source is not available in this task, say so explicitly and explain how that limits the conclusion.

Evidence gaps are not failures by themselves. A gap becomes dangerous when the final deliverable hides it. For example, a spreadsheet comparing vendors should not score “security posture” as if all vendors provided comparable security documentation when only one vendor’s current documentation was available. The acceptance standard should require either a missing-data marker or a confidence downgrade.

Evidence issue Safe treatment in the artifact Unsafe treatment to reject
Unavailable connected source List the source as unavailable and describe the decision impact. Infer its contents from prior knowledge or similar sources.
Conflicting source dates Use the current source for present-state claims and older sources only for history. Blend old and new facts into one uncited statement.
Secondary commentary without primary support Label it as interpretation and avoid using it as the sole basis for a recommendation. Present commentary as verified fact.
Missing stakeholder perspective Add a follow-up question, interview need, or review gate. Claim consensus without checking the missing group.

Turn the research into the right editable artifact

OpenAI says Deep Research in Work and Codex can turn findings into an editable document with citations, and users can request a document, presentation, spreadsheet, or Site when the needed tools are available. Phrase the request as an output contract, not as a formatting preference: specify the artifact type, audience, sections, citation treatment, review gates, and what must remain editable.

Artifact request: editable report

Convert the research into an editable report for [audience].

Requirements:
1. Include an executive summary, decision context, evidence table, analysis, recommendation, risks, evidence gaps, and appendix.
2. Keep citations or source links attached to the claims they support.
3. Label evidence, inference, uncertainty, and unsupported items separately.
4. Use the supplied template or structure if available in this task.
5. Do not remove contradictions; summarize how they affect confidence.
6. After creating the artifact, provide the file or link and a short verification checklist I should complete before sharing.

For presentations, ask for slide-level claims and speaker notes rather than a pasted report broken into slides. A slide deck should compress evidence without stripping traceability, so require each slide to include either source notes, speaker-note citations, or an appendix mapping claims to sources.

Artifact request: editable presentation

Create an editable presentation for [audience] and [meeting purpose].

Structure:
- Slide 1: Decision to be made
- Slide 2: What changed or why now
- Slide 3: Evidence summary
- Slide 4: Options considered
- Slide 5: Recommendation
- Slide 6: Risks and mitigations
- Slide 7: Evidence gaps and follow-up owners
- Appendix: claim-to-source mapping

Keep the deck concise, but preserve citations in speaker notes or appendix form. Do not include sensitive source excerpts unless they are necessary and permitted for this audience.

For spreadsheets, force the system to define columns before generating rows. This reduces the chance that the spreadsheet becomes a flat dump of findings with inconsistent scoring. Ask for source columns, confidence levels, missing-data markers, and scoring definitions so a reviewer can audit each row.

Artifact request: editable spreadsheet

Create an editable spreadsheet from the research.

Before finalizing, define the columns and scoring rules. Include:
- Item or vendor
- Category
- Evidence-backed finding
- Source or citation
- Confidence level
- Missing evidence
- Score, if scoring is justified
- Rationale for score
- Human review required: yes/no

If comparable evidence is missing for a row, do not assign a confident score. Use "insufficient evidence" and explain what source would be needed.

If you request a Site, keep the same discipline. A Site can make research feel publication-ready, but OpenAI’s guidance still requires the user to verify the output file or link, inspect citations, remove sensitive material, and confirm rights and permissions before external sharing. Do not treat Site generation as autonomous publication or as proof that the material is approved for a public audience.

Artifact request: editable Site, when supported

If the required tools and permissions are available, create an editable Site for [internal/external audience].

Requirements:
1. Organize the Site into decision context, findings, recommendation, risks, evidence gaps, and sources.
2. Keep source links or citations visible enough for reviewers to audit claims.
3. Do not publish or distribute externally without my explicit review.
4. Flag any sensitive, licensed, confidential, or long-quoted material that may need removal or permission review.
5. Provide the Site link and a checklist for validating access, editability, citations, and audience suitability.

Apply artifact acceptance checks before you reuse the output

The final step is not “read the answer”; it is acceptance testing. OpenAI’s documentation tells users to verify underlying sources, especially for decisions, external publishing, or high-stakes topics, and to verify that generated files or links exist and open successfully. Build that instruction into a repeatable acceptance checklist for every artifact.

Acceptance check How to perform it Reject or revise if
Artifact opens Open the generated document, presentation, spreadsheet, or Site link in the intended workspace context. The file is missing, inaccessible, view-only when editing was required, or visible to the wrong audience.
Citations are usable Spot-check citations for the highest-impact claims and verify that each cited source supports the statement. A citation points to a source that is irrelevant, outdated, inaccessible to reviewers, or weaker than the claim.
Evidence boundaries are preserved Compare the artifact against the original scope, allowed sources, exclusions, and timeframe. The artifact includes out-of-scope claims without labeling them as background or uncertainty.
Contradictions are visible Review the contradictions section or appendix and confirm that unresolved conflicts affect confidence or recommendations. The artifact buries contradictory evidence or presents a contested conclusion as settled.
Evidence gaps are explicit Check that missing sources, unavailable connected apps, and unsupported assumptions are listed with follow-up actions. The artifact replaces missing evidence with confident language or invented certainty.
Sensitivity and rights are reviewed Remove confidential details not needed for the audience and check whether quoted or included material can be reused. The artifact contains sensitive information, excessive copied passages, or material whose sharing rights are unclear.

For enterprise and education environments, add an access-control check before circulation. Deep Research respects the sources and permissions available in the task, but the finished artifact may be easier to forward than the original sources. Confirm that viewers are allowed to see the cited material, embedded excerpts, uploaded-file content, and connected-app-derived findings before sending the artifact outside the original review group.

Use a final revision prompt to force traceability before handoff

After you inspect the artifact, send one final revision prompt that converts your acceptance findings into specific edits. This is more reliable than asking for a generic polish pass, because polish can remove caveats, compress citations, or soften uncertainty language that reviewers need.

Final revision prompt: acceptance-driven cleanup

Revise the artifact using these acceptance findings:

1. Claims needing stronger citation: [list]
2. Claims to downgrade or mark uncertain: [list]
3. Sources that reviewers cannot access: [list]
4. Sensitive or rights-sensitive material to remove or summarize: [list]
5. Missing evidence to keep visible: [list]
6. Sections requiring clearer decision language: [list]

Do not remove citations, contradiction notes, confidence labels, or evidence-gap tables. Make the artifact more decision-ready without making the conclusions sound more certain than the evidence supports.

The operational rule is simple: steer early, audit before synthesis, and accept the artifact only after it opens, remains editable where required, preserves citations, and exposes uncertainty. Deep Research is most valuable when it accelerates the route from sources to a reviewable deliverable, not when it replaces the human responsibility to verify evidence, permissions, and audience suitability.

Verification and operations: turn a polished research artifact into a defensible work product

A Deep Research result should be treated as a drafted work product, not as a completed decision record. OpenAI’s Deep Research guidance says users should verify underlying sources, especially for decisions, external publishing, or high-stakes topics, and it explicitly warns that citations improve traceability but do not guarantee correctness. The operational goal after the run is therefore to separate four things: what the model claimed, what the cited source actually says, whether the cited source is authoritative for the decision, and whether the generated file can be opened, edited, shared, and retained under your organization’s rules.

Use the verification phase even when the report looks well structured. A report can contain real citations but still overstate a finding, cite a source that supports only part of the sentence, miss a newer contradictory source, mix evidence with inference, or produce an editable file that does not open in the target tool. The safest workflow is to make verification visible: create a claim-to-source audit table, sample citations by risk, run a sensitivity and rights review, open every deliverable, check allowance consumption, and request a final revision that incorporates only verified corrections.

Build a claim-to-source audit before approving the answer

A claim-to-source audit converts the final narrative into a reviewable evidence ledger. Instead of asking whether the report “has citations,” the reviewer asks whether each material claim is supported by the cited source and whether the claim’s wording matches the strength of the evidence. This matters because Deep Research can produce citations or source links, but OpenAI’s documentation does not say citations are proof of correctness or source authority.

Audit field What to record Decision rule
Claim ID Assign a short identifier such as C-01, C-02, or RISK-03. Every recommendation, number, comparison, legal/regulatory statement, vendor assertion, and strategic conclusion gets an ID.
Generated claim Copy the exact sentence or table cell from the report. Do not paraphrase during audit; wording strength is part of the risk.
Cited source Record the citation or source link supplied by Deep Research. If no citation exists for a material claim, mark it unsupported until fixed.
Source says Summarize what the source actually supports after human inspection. Distinguish exact support, partial support, contradiction, stale information, and no support.
Reviewer action Approve, soften wording, add caveat, replace source, add missing source, or remove claim. Any high-impact decision claim must be approved by a responsible human reviewer before reuse.

For developer and research teams, store this audit beside the deliverable rather than inside a transient chat only. For enterprise administrators, require the audit table for externally shared market maps, vendor scorecards, due-diligence summaries, policy briefs, security assessments, customer research syntheses, and executive memos. OpenAI Academy recommends validating key claims and citations, and these artifacts are exactly the kinds of documents where a polished but weakly supported conclusion can travel quickly.

Sample prompt: ask Deep Research to produce an audit-ready table

Prepare a claim-to-source audit table for the report you just produced.

For each material claim, include:
1. Claim ID
2. Exact claim text
3. Citation or source link
4. Evidence type: primary source, vendor source, news source, internal file, expert interpretation, or unsupported
5. Whether the source directly supports, partially supports, contradicts, or does not support the claim
6. Confidence level and reason
7. Recommended human review action

Do not add new claims. Mark any unsupported or weakly supported claim plainly.

Use citation sampling by risk, not by convenience

A practical citation audit does not always require checking every footnote with the same intensity. It does require checking the citations that can change a decision. Use a risk-weighted sample: verify all citations for high-impact claims, then sample medium- and low-impact claims enough to detect systematic problems. If the first sample reveals missing, irrelevant, or overstated citations, expand the audit because citation quality problems often cluster by source type or section.

  1. Verify 100% of decision-critical claims. Check every citation behind recommendations, budget assumptions, compliance statements, security implications, customer commitments, competitive comparisons, and quantified findings.
  2. Verify 100% of surprising or counterintuitive claims. A claim that contradicts common assumptions may be valuable, but it should not pass review on a citation label alone.
  3. Verify all citations drawn from user files or connected apps that contain sensitive internal context. Starting Deep Research does not grant new access, according to OpenAI, but the output can still expose information from authorized sources if the prompt allowed it.
  4. Sample at least several routine background claims from each major section. If the model consistently cites relevant sources for low-risk background, keep the sample bounded; if it does not, escalate.
  5. Check recency where time matters. For market, product, legal, security, and pricing topics, a correct older source may still be the wrong basis for a current recommendation.

The sampling result should be explicit. Mark the document as “citation-sampled,” “fully citation-checked,” or “not citation-checked,” and state which categories were reviewed. This prevents an executive summary from implying more verification than actually occurred.

Run sensitivity, confidentiality, and rights review before sharing

OpenAI Academy advises users who plan external sharing to recheck citations, remove sensitive information, confirm rights and permissions for included material, avoid long copied passages, and retain citations. Turn that guidance into a release gate. The release gate should ask whether the report includes confidential strategy, customer identifiers, employee information, unreleased financial or product plans, licensed third-party material, long copied passages, or material from connected apps that was authorized for research but not authorized for redistribution.

Review area Questions to ask Safe action
Sensitive business information Does the output reveal internal plans, pricing strategy, vendor negotiations, incident details, or nonpublic metrics? Redact, aggregate, or produce an internal-only version.
Personal data Does the artifact include names, emails, transcripts, customer records, or interview details that are not needed? Remove identifiers unless a lawful and approved business purpose requires them.
Third-party rights Does the report reproduce long passages, proprietary tables, images, or paid research content? Replace long copied material with short summaries and preserve citations.
Connected-source boundaries Was the source accessible for internal analysis but not approved for external distribution? Create a public-safe version that cites public sources or omits restricted material.

For enterprise administrators, this gate should be aligned with workspace role-based access, provider permissions, and any internal classification scheme. OpenAI states that Deep Research uses available and authorized sources in the current task and does not grant additional app or file access; however, authorized access for research is not the same as permission to republish the resulting synthesis.

Open every generated file, link, spreadsheet, presentation, or Site

OpenAI’s Deep Research documentation says editable reports, presentations, spreadsheets, or other files depend on the task and supported capabilities, and users must verify that the output file or link exists and opens successfully. This is an operational requirement, not a cosmetic check. A file that fails to open, opens with broken formatting, loses citations, or cannot be edited in the expected tool is not an accepted deliverable.

  1. Open the artifact from the same account and workspace that will use it. This catches permission problems that are invisible inside the original chat.
  2. Confirm the artifact type. If you requested a spreadsheet, verify that formulas, tabs, column headings, and source references survived export or creation.
  3. Check citations in the artifact, not only in the chat. Citation links, footnotes, and source tables can be lost or reformatted when moving from report to presentation or spreadsheet.
  4. Test editability. Make a small copy edit, add a comment, or change a cell to verify that the deliverable is not merely a static preview.
  5. Check sharing behavior before sending. Confirm that intended recipients can open the item and unintended recipients cannot, using your organization’s approved access process.

If the artifact does not open, ask for a replacement format rather than trying to reconstruct the missing work manually. If the content opens but lacks traceability, request a version that includes a source appendix, claim IDs, or citations per table row.

Monitor Work and Codex allowance while research is running

Deep Research in Work uses the existing Work/Codex allowance or credits, and OpenAI says it does not consume the separate Chat Deep Research task allowance. The practical implication is that a long research run can compete with coding, local work, or other Work/Codex tasks that depend on the same pool. Teams should treat research scope as a resource decision, especially when asking for multiple audience versions, spreadsheets, presentations, or repeated revisions.

Set an allowance checkpoint before launch. Define the maximum number of research passes, the expected artifact count, and the point at which the user should stop the run and narrow scope. OpenAI’s Work and Codex guidance also notes that task consumption can vary by task, input and output size, reasoning settings, and speed choices. Do not assume that two similarly titled research tasks will consume the same amount of allowance.

Account for storage, training, and local-versus-cloud expectations

Retention and data-use expectations should be reviewed before uploading source files or authorizing connected apps. OpenAI’s Deep Research guidance states that Business, Enterprise, and Edu content is not used for training by default, while Plus and Pro conversations may be used unless training is disabled in Data Controls. That distinction is important for founders, consultants, and independent researchers who may use personal or professional accounts differently across projects.

Work and Codex also require careful local-versus-cloud assumptions. OpenAI’s Work and Codex guide says Work on web and mobile runs in the cloud, while desktop Work can use local files and apps with permission; it also notes that messages and task context may still be stored in the cloud even when work runs locally. Do not tell employees that local file access means “nothing leaves the device” unless your administrator and current product documentation support that statement for the specific workflow.

Close the loop with a verified revision pass

The final revision should be driven by audit findings, not by a vague request to “make it better.” Give Deep Research the corrections, unsupported claims, source replacements, sensitivity edits, and artifact defects in a structured list. Ask it to revise without introducing new uncited claims, and require it to preserve uncertainty where the evidence remains incomplete.

Revise the deliverable using only the verification notes below.

Rules:
- Remove or soften claims marked unsupported or partially supported.
- Preserve citations for every material claim.
- Add an "Evidence gaps and human review required" section.
- Do not introduce new facts unless you cite them.
- Remove sensitive or rights-restricted material identified in the review.
- Keep the requested format and repair any broken tables or source references.

Verification notes:
[Paste claim-to-source audit findings, sensitivity review findings, and file-open issues]

Reusable operating checklist for Deep Research in Work and Codex

  • Define the decision. State the audience, timeframe, output format, evidence boundary, and review standard before the run.
  • Confirm source authorization. Use only files, web sources, and connected apps that are permitted for the task; starting Deep Research grants no new access.
  • Steer while live. Interrupt when the plan is too broad, the sources are weak, or the report is drifting toward unsupported synthesis.
  • Request traceability. Ask for claim IDs, citations, source tables, contradiction notes, and evidence gaps.
  • Audit citations. Remember that citations do not prove correctness; check whether each cited source actually supports the claim.
  • Review sensitivity and rights. Remove confidential, personal, restricted, or over-copied material before sharing.
  • Open every artifact. Verify that each document, presentation, spreadsheet, or Site exists, opens, remains editable, and retains citations.
  • Watch allowance impact. Work Deep Research uses Work/Codex allowance or credits, not the separate Chat task allowance.
  • Apply account and workspace rules. Consider web search enablement, RBAC, plan, region, connected-provider permissions, and data controls.
  • Revise from evidence. Feed verified corrections back into the final pass and preserve unresolved uncertainty.

The disciplined way to use Deep Research in Work and Codex is to make the model fast at gathering, organizing, and drafting while keeping humans responsible for evidence quality, permissions, and release decisions. When source boundaries, live steering, citation audits, artifact verification, and revision gates are part of the workflow, Deep Research becomes a practical research-production system rather than a polished shortcut around review.

Access 40,000+ AI Prompts for ChatGPT, Claude & Codex — Free!

Subscribe to get instant access to our complete Notion Prompt Library — the largest curated collection of prompts for ChatGPT, Claude, OpenAI Codex, and other leading AI models. Optimized for real-world workflows across coding, research, content creation, and business.

Get Free Access Now →

Useful Links

Get Free Access to 40,000+ AI Prompts for ChatGPT, Claude & Codex

Subscribe for instant access to the largest curated Notion Prompt Library for AI workflows.

More on this