Deep research runs are metered separately from ordinary messages, and the number you get depends on your plan. OpenAI documents current allowances on its plan pages rather than in the API reference, and those figures change often enough that checking directly is more reliable than trusting a number quoted anywhere else, including here.

What is worth understanding instead is why this feature is rationed at all, what consumes a run, and how to avoid burning an allowance on a question that did not need it.

Key points

  • Deep research is metered separately from normal chat messages, on a per-plan allowance.
  • Each run is expensive to produce, involving many searches, page reads and a long reasoning process.
  • The allowance resets on a cycle tied to your plan rather than accumulating indefinitely.
  • Most questions do not need it. Ordinary chat with web search answers a large share of what people spend runs on.
  • Prompt quality decides the result. A vague brief produces a long document that answers something adjacent to your question.

Why deep research is rationed when normal messages are not

The reason is straightforward once you look at what a run actually does.

A standard chat response is one pass through a model. A deep research run is a sustained process: it plans an approach, issues many searches, opens and reads a large number of pages, evaluates what it found, follows up on gaps, and then writes a long structured document with citations. That can involve hundreds of retrievals and a great deal of reasoning before a single word of output appears.

In token terms a single run can consume more than a week of ordinary conversation. That is why it sits behind its own allowance rather than counting against a general message limit, and it is also why the allowance is small in absolute terms.

The practical consequence is that runs should be treated as a scarce resource with real value, not as a better default for any question that feels substantial.

How the allowance behaves in practice

Two properties of the metering matter more than the specific number, and both are stable even as the figures change.

Allowances reset on a cycle rather than accumulating. Unused runs do not roll over, which means a month where you needed none does not fund a month where you need many. That argues for using the allowance rather than hoarding it, but using it on questions that fit rather than on whatever arrives first.

A failed or unsatisfying run generally still counts. If the brief was vague and the output missed the point, the budget is spent regardless. This is the single strongest argument for spending five minutes on the brief, because the marginal cost of a better question is nothing and the marginal cost of a wasted run is a meaningful fraction of the period’s allowance.

There is also a practical consequence for teams. Because the allowance sits on an individual account, a team that relies heavily on this feature ends up rationing across people rather than across questions, and the person who happens to have runs left is not necessarily the person with the best question. Teams that use it seriously tend to agree in advance what qualifies.

What a run is genuinely good for

Deep research earns its cost on a specific shape of question, and recognising that shape is most of what separates people who find it valuable from people who find it disappointing.

It works well when the answer requires assembling information from many sources that no single page contains. Comparing how a dozen vendors price a category. Establishing what several jurisdictions require on the same regulatory point. Tracking how a technical standard changed across successive revisions. Building a landscape of who is doing what in a field.

It works poorly when a single authoritative source already holds the answer. A vendor’s own documentation, a statistical agency’s published table, a company’s filings. In those cases ordinary chat with search finds it in seconds and a deep research run spends its budget confirming something that was never in doubt.

It also works poorly on questions with no stable answer. Anything highly current, contested, or dependent on information that is not published will produce a long confident document assembled from whatever was available. Length is not evidence, and a thorough-looking report built on thin sources is more dangerous than a short answer that admits uncertainty.

The distinction worth carrying is between questions where the information exists but is scattered, and questions where the information does not exist. Deep research is excellent at the first and produces its most misleading output on the second, because it will assemble something regardless.

What to know before deciding to spend a run

Several things affect whether you get value, and most of them are decided before the run starts.

The brief determines the output. A vague question produces a broad survey. A specific question with stated scope, timeframe and what you intend to do with the answer produces something usable. This is the single largest factor.

It cannot see what is not published. Paywalled material, internal documents and anything behind a login are outside its reach unless you supply them.

Citations need checking. The output includes sources, which is the point, but a citation supporting a claim loosely is a real failure mode. The value of the citation is that it lets you verify quickly, not that verification becomes unnecessary.

Runs take time. Expect minutes rather than seconds, which changes how you use the feature. Starting a run and getting on with something else is the sensible pattern, and waiting for it wastes more of your time than the run costs.

Recency has limits. Search finds what is indexed. For events of the last few hours, ordinary search with a targeted query is frequently better.

Writing a brief that does not waste the run

  • State the decision behind the question. “I am choosing between these approaches for a team of ten” produces something different from “tell me about these approaches”.
  • Give scope explicitly. Which markets, which time period, which kinds of source you trust and which you do not.
  • Name what you already know. This stops the run spending its budget re-establishing your starting point.
  • Say what output shape you want. A comparison table, a chronology, a set of options with trade-offs.
  • Exclude what you do not want. Ruling out categories saves real budget and sharpens the result.
  • Ask for uncertainty to be marked. Requesting that unverified or thinly sourced claims be flagged is the most useful instruction most people never give, and it costs nothing to include.
  • Set a length expectation. Without one you tend to get maximum length, and a shorter, denser answer is often more useful than a long one you will skim.

What deep research costs to run, and why that shapes the limit

It helps to see the underlying economics, because they explain why the allowance is shaped the way it is and are unlikely to change soon.

Published API pricing gives a sense of scale. GPT-6 Astra is billed at $10 per million input tokens and $50 per million output, with reasoning tokens counted as output. A research run that reads a hundred pages pulls a very large amount of text into context, and it reasons extensively across that material before writing.

Long inputs are also repriced. OpenAI documents that requests above 272,000 input tokens are billed at double the input rate for the whole request. A run that accumulates a substantial body of source material crosses that boundary comfortably, which means the cost does not rise linearly with the amount read.

None of this is visible to someone using ChatGPT on a subscription, where the cost is absorbed into the plan. But it explains why the feature carries a separate allowance instead of counting as a few ordinary messages, and why that allowance is measured in a handful of runs rather than hundreds.

It also suggests where the value is. A run that saves an afternoon of manual searching across many sources is worth its cost several times over. A run that confirms something a single documentation page already stated is not, and the difference is entirely in the question you asked.

Decision framework

Five questions before starting a run.

  1. Would one authoritative source answer this? If yes, find that source instead. Vendor documentation and official statistics do not need a research run.
  2. Does the answer require synthesis across many sources? This is the case that justifies the cost.
  3. Can I state what decision this informs? If not, the brief will be vague and so will the output.
  4. Is the information likely to be published at all? Runs cannot find what nobody wrote down.
  5. Do I have time to verify it? An unverified research document is a liability rather than an asset, particularly one that reads as authoritatively as these do.

Knowing which tool fits which question, and how to brief it so the output is checkable rather than merely plausible, is a skill rather than a product feature. If you want a structured route in, explore Coursiv AI lessons and check current plan details on the official site.

Your next step

Before your next run, write the brief in three sentences: what decision you are making, what you already know, and what shape of answer would be useful. Then read it back and ask whether a single well-chosen search would answer it.

A surprising share of the time it would, and recognising that is worth more than any technique for using the feature better. The runs you save go to questions that genuinely need many sources assembled, which is where the allowance pays for itself.

FAQ

How many deep research runs do I get?
The allowance depends on your plan, and OpenAI updates these figures periodically as capacity and pricing change. Check the current plan details on the official site rather than relying on a number quoted in an article, including this one.
Why is deep research limited when normal messages are not?
A single run performs many searches, reads a large number of pages and reasons at length before writing. It consumes substantially more compute than ordinary conversation, which is why it carries its own allowance.
When does my allowance reset?
Resets follow a cycle tied to your plan. The plan page is the reliable place to check, since both the allowance and the cycle have changed over time.
Is deep research more accurate than normal chat?
It is better sourced, not automatically more accurate. It cites what it used, which makes verification much faster, but a citation can support a claim only loosely and still look convincing on the page. Treat the sources as a checking aid rather than as proof that checking already happened.