TL;DR

  • ChatGPT has a feature called Fast answers.1 It appears to serve ready-made answers to frequently asked, information-seeking questions.
  • What if I were the person responsible for launching it? For fun, I worked backward from the product and wrote a hypothetical system design.

This is a four-part series. Part 0 covers my understanding of the product and a survey of the related literature. The rest of the system design continues in the following posts:
Part 1. Business Context and ML System Design
Part 2. Production Architecture
Part 3. Operations and Continuous Improvement

1. ChatGPT Does Not Always Generate a New Response

One day, I was solving a LeetCode problem and wanted to refresh my memory on quickselect. ChatGPT has become my go-to knowledge hub, so I naturally typed quickselect into it.

This time, though, the answer came back surprisingly fast and was labeled Fast answer.

A Fast answer for the query "quickselect"

Figure 1. My first encounter with a Fast answer in ChatGPT.

It was the first time I had seen the feature, and I was curious enough to try the same query again. I opened another chat, typed quickselect, and got the same answer.

That was unusual. An LLM will generally produce somewhat different responses even when it receives the same prompt twice. This time, the wording, structure, and examples were identical. My immediate guess was that ChatGPT was returning a stored query-answer pair instead of generating a fresh response. Query-answer caching would be a reasonable way to reduce both latency and inference cost for common questions with stable answers.

That led me to a more interesting question: if I were building and launching this feature, how would I design it?

What follows is a hypothetical design based on public information, a few product experiments, and assumptions I made myself. I have no inside knowledge of how Fast answers is actually implemented.

2. Product Observation

A quick search did not reveal much, but the feature does appear in the ChatGPT release notes.1

The April 22, 2026 ChatGPT release note for Fast answers

Figure 2. OpenAI’s release note introducing Fast answers.

The note says Fast answers are intended for common information-seeking questions, such as “Show me the Seven Wonders of the World” or “Which football team has the most Super Bowl titles?”2 ChatGPT may respond faster when two conditions hold:

  1. The question does not require a personalized response.
  2. ChatGPT has a high-confidence answer ready.

It also says that Fast answers do not reference the user’s past chats or memory.

The release note does not explain what “has a high-confidence answer ready” means technically. My working hypothesis is that ChatGPT maintains a collection of pre-generated answers for common, straightforward questions and returns one when an incoming query matches with sufficient confidence. That is the assumption I use throughout this series.

A few small experiments helped clarify the product surface before considering the system behind it. These were informal observations rather than a controlled evaluation, and the behavior may change as the feature evolves.

2.1 The Same Query Returns the Same Answer

Several frequently searched topics triggered Fast answers, including:

  • quickselect
  • why is the sky blue
  • what is ROAS

For these experiments, Fast answers were enabled under Settings → Personalization, with Intelligence set to Instant. With this setup, each query produced a Fast answer. Repeating a query in a new chat returned the same response.

Fast answer for why is the sky blue Fast answer for what is ROAS

Figure 3. Fast answers for why is the sky blue and what is ROAS.

Below a Fast answer, ChatGPT shows a lightning-bolt icon. Clicking it lets the user regenerate the answer immediately or describe in natural language how it should be regenerated.

This button is probably one of the main channels for collecting user feedback. The system design later in this series will discuss how feedback from this surface could be used.

The menu behind the lightning-bolt icon on a Fast answer

Figure 4. The Fast answer menu offers regeneration and a free-form instruction field.

2.2 The Same Query Does Not Always Trigger It

Even the exact same query did not always trigger a Fast answer. In the screenshot below, ChatGPT generated a response to quickselect and began answering in Korean. I normally use ChatGPT in Korean, so my memory or conversation history was probably included in the context for this response.

A normally generated Korean response for the query "quickselect"

Figure 5. The same quickselect query did not trigger a Fast answer this time.

There are several possible reasons why the feature triggers inconsistently. OpenAI may still be running conversation-level A/B tests. There may be a heuristic that treats repeated submissions of the same query as dissatisfaction with the Fast answer and switches back to generation. Or some hidden logic may decide whether to show a Fast answer for each user based on personalized signals. This last possibility will be discussed in more detail in the system design.

Regardless of the reason, randomized triggering may actually be the better launch strategy. Rather than always showing the cached result for the same query, it may be better to randomize whether the user receives the cached answer. The resulting data—including user reactions and inference-cost savings—could inform the next product decision. Randomized data would also reduce bias in later analysis and make counterfactual analysis more credible. The system design will cover this in more detail.

2.3 Semantically Similar Queries Return Different Answers

At first, I thought the feature might use semantic search to retrieve answers for similar queries. For example, quickselect and what is quickselect can both be satisfied by a single answer explaining the concept. It would be reasonable to prepare one answer and return it for semantically similar user queries.

Fast answer for quickselect A different Fast answer for what is quickselect

Figure 6. quickselect and what is quickselect trigger different Fast answers.

The actual behavior was different. As the screenshots show, even semantically equivalent queries returned different answers. This suggests that the current Fast answer system may store one answer per query and determine cache hits through exact query matching.

There are two tradeoffs to consider:

  1. Maintain one polished answer, which makes answer-quality management easier, but requires solving semantic cache hit and miss decisions accurately.
  2. Use exact matching for queries, which keeps the lookup logic simple but increases the number of query-answer pairs to manage.

From the perspective of launching and operating an initial system, the second approach may have been simpler. It makes sense if the product decision is that user disappointment from a false cache hit and a low-quality answer is more costly than the LLM generation cost of a false cache miss. The system design will examine this tradeoff in more detail.

2.4 It Appears to Support Multiple Languages

The feature also triggered outside English, suggesting that eligibility is not limited to one language. As in Section 2.3, two questions with similar meanings received different answers. The screenshots below ask for the definition of ROAS using two different Korean expressions, and the responses are different as well. This supports the exact-query-matching hypothesis regardless of language.

Korean Fast answer for ROAS가 뭐야? Korean Fast answer for ROAS 뜻?

Figure 7. Two Korean queries asking for the definition of ROAS return different Fast answers.

This was interesting because limiting an initial launch to English would normally be safer. Cached answers need to be reviewed for quality before they are served. Multilingual review adds outsourcing, coordination, and operational costs, so supporting a language with relatively few global speakers, such as Korean, would not be easy.3

This suggests two possibilities. First, OpenAI may believe that the cost savings from multilingual Fast answers are large enough to justify the additional review cost. Second, it may have enough confidence in an automated quality-review system to operate the feature across languages.

2.5 Fast Answers Have a Faster TTFT

Browser developer tools showed a clear difference in how quickly the two response types began arriving. Using time to first response byte as a proxy for TTFT, Fast answers took roughly 400–600 ms, while normally generated answers took roughly 1.8–2.2 seconds.

Network timing for a Fast answer

Figure 8. The Fast answer began arriving after about 559 ms.

Network timing for a normally generated answer

Figure 9. The normally generated answer began arriving after about 1.81 seconds.

Interestingly, Fast answers were still delivered as a stream, even though the answers appeared to have been prepared in advance. These screenshots were taken in July 2026. When the feature first launched two months earlier, I remember the entire answer arriving more quickly without visible streaming. From a UX perspective, the response appeared to fill in immediately, which made the feature feel extremely fast.

Something may have changed in the meantime. There are at least three possible explanations:

  1. The pre-generated-answer hypothesis is wrong, and Fast answers are actually generated by a very lightweight model.
  2. The answers are pre-generated, but OpenAI switched to streaming to keep the UX consistent with normal responses and may be A/B testing that presentation.
  3. The behavior changed for some other reason that cannot be observed from the outside.

The goal of this series is not to reverse-engineer the feature exactly. It is to develop my own system design from the product behavior I can observe. I will therefore continue with the assumption that Fast answers return pre-generated responses.

2.6 It May Trigger in the Middle of a Conversation

I expected this feature to apply only to the first query in a conversation. According to OpenAI’s study of how people use ChatGPT, 24% of messages were classified as Seeking Information as of July 2025—in other words, use that closely resembles search.4 Information-seeking traffic should contain frequently repeated queries, so I thought this feature was introduced to reduce LLM inference cost by caching answers to those queries.

However, a comment in a Reddit discussion about Fast answers complains that a keyword triggered a Fast answer in the middle of a multi-turn conversation. In that case, ChatGPT ignored the context accumulated so far and returned an answer that had already been prepared.

That behavior was unexpected. I do not know whether it is intended or a bug. In my design, I would probably allow Fast answers only for the first query. Once a conversation becomes multi-turn, the user is more likely to care about the context built up so far, which makes it much harder for a context-free Fast answer to satisfy them.

3.1 Direct Answers Before LLMs

Fast answers may look like a new LLM product feature, but search engines have long faced a similar problem. Google Featured Snippets extract an answer-like passage from a webpage and place it above the ordinary search results.

A Google Featured Snippet answering the query "Why is the sky blue" Figure 10. A Google Featured Snippet answering the query “Why is the sky blue.”

The difference between Google and ChatGPT is that Google has to select good answers from external content in advance, while ChatGPT has to select good answers from responses generated internally.

Google’s blog post suggests three insights.

First, Google must have invested substantial human-rater effort in evaluating search quality. The post discusses cases in which quality became a problem and shares a 182-page document called the Search Quality Rater Guidelines. I expect that ChatGPT made a similar effort to ensure that Fast answers are trustworthy.

Second, retrieval and the serving decision are separate problems. Storing and retrieving a good snippet is one problem; deciding whether to show it to the user is another. Google did not show a snippet when its authority, quality, or compatibility with the query was insufficient. Similarly, I will describe Fast answers as a three-stage system: curation, retrieval, and policy.

Third, Google worked on cases in which queries were lexically similar but had different semantic intent. As mentioned above, its approach may have been to keep the number of snippets small to reduce management costs while solving the semantic query-matching problem. ChatGPT, however, provides internally generated content, which should make that content easier to generate and manage than Google’s external content. With that in mind, Fast answers may not have needed semantic query matching in its initial version.

3.2 Managing QA Catalog

While I could not find prior work that exactly matches the Fast answers setting, there is a substantial body of work on automatically building and maintaining FAQ-style knowledge bases from historical user interactions.5

The common idea is to mine recurring information needs from historical user queries or conversations, group semantically similar questions, and turn them into reusable question-answer pairs. These pairs can then form a prepared-answer catalog that a separate retrieval system may consult when a similar request arrives in the future.

One particularly relevant example is AI Knowledge Assist. This paper follows the common approach described above and provides a concrete example of how such a system can be built.

Overview of the AI Knowledge Assist pipeline Figure 11. Overview of AI Knowledge Assist. Source: Figure 2 in Laskar et al. (2025), licensed under CC BY 4.0.

The paper assumes that many companies want to build conversational AI chatbots or RAG systems but do not have company-specific knowledge bases. However, if a company has customer service chat data, QA data is already embedded in the agents’ answers to customer questions. The paper therefore proposes a way to extract and clean this data and keep it updated over time. This is similar to ChatGPT’s situation: user chat logs already exist, and the task is to extract a QA set from them. However, the same question may have different answers in ChatGPT depending on each user’s context, so those answers cannot be used directly. The data could at least be used to identify recurring questions.

The paper uses LLMs in most stages. It first uses an LLM to extract reusable QA pairs from raw conversation data. Rather than extracting arbitrary pairs, it applies the following conditions: information-seeking, non-personalized, no PII, not time-sensitive, universal, and useful. It then clusters the QA pairs by question similarity and uses an LLM to select a representative QA pair from each cluster. A pair is either added automatically or sent for human review. The paper also mentions an automatic update mechanism that uses question similarity and answer similarity for future catalog updates.

In my view, there are three major points to consider when building a catalog system.

First, QA pair filtering has to be done well. It effectively determines the final quality of the product. A clear policy is needed to define which kinds of questions align with the product’s goals. This is not something that can be solved simply by having a product owner write down a constitution-like set of rules. The initial policy will most likely contain many vague areas, creating conflicts no matter how closely a human rater or an LLM rater tries to follow it. The policy therefore has to improve through multiple iterations. A previous post from my time at Hyperconnect discusses this process in more detail.6

Second, the decision to add a QA pair to the catalog will also likely be automated with an ML model, but designing that decision policy is not easy. For example, when setting a confidence cutoff, several variables have to be considered: whether the confidence comes from a calibrated model, how the cutoff affects human-rater costs, and how much total traffic the system receives. The problem is to find a Pareto-optimal point among these variables.

Third, the system has to support continuous updates. Even seemingly stable catalogs can become stale over time, and the distribution of what counts as a good QA pair is likely to shift. The system should therefore be designed from the beginning to be robust to this shift. For example, as user-feedback data accumulates, the system should automatically adapt to distribution shifts so that its performance is maintained or improves over time.

3.3 Retrieval

Once a good QA catalog has been built, the system has to retrieve the right entry for each user query. Although this system design assumes exact query matching to reduce system complexity, it is still worth identifying and comparing the other available options.

Information retrieval in modern ML systems is already a well-studied problem. Several papers have also explored it specifically in the FAQ retrieval domain.7 Each presents a novel idea, but the core modeling decisions come down to two questions:

  1. How should relevance between a user query and a QA pair be defined?
  2. Should the retrieval target consider only the question, or the answer as well?

As with any ML modeling problem, defining the objective is the most important step. For Fast answers, the retrieval objective could be aligned directly with the product objective: maximizing user satisfaction with the answer. However, user satisfaction can be sparse, noisy, and difficult to measure. In that case, an LLM-based prediction of user satisfaction could be used instead. An answerability score from a domain expert could also serve as a training proxy. Whatever signal is chosen, it should be quantifiable and aligned with the product objective, and the model should be trained to optimize it.

Query pair Answer Cosine similarity
When was Python created? ↔ Who created Python? 1991 vs. Guido van Rossum 0.927
Who created Python? ↔ Who was the original author of Python? Both Guido van Rossum 0.853

Table 1. Cosine similarities measured with BAAI/bge-small-en-v1.5.

The important point is that an off-the-shelf pretrained sentence embedding is unlikely to be aligned with the Fast answers objective. A general sentence embedding model considers the first query pair more similar than the second. However, the first pair requires different answers, while the second pair shares the same answer. Query similarity alone is therefore not necessarily aligned with answer equivalence. Fine-tuning for the actual objective is likely to be essential.

3.4 Final Decision Layer

So far, the discussion has focused on which QA pairs should be included in the candidate set and how to select the best one among them. A production-level system has to consider one more step: the decision policy that determines whether the selected QA pair should actually be shown to the user. In other words, the system has to answer this question automatically: For this user, in this situation, is showing a Fast answer better than normal generation? The need for this layer will be covered in the next system design post. For now, it is useful to examine how the QA domain has approached this problem.

vCache: Verified Semantic Prompt Caching8 is the paper most directly related to Fast answers. When a user query and a cached query are similar in embedding space, the cached response is reused. Rather than using one fixed similarity threshold, vCache assigns a separate threshold to each cached entry and learns those thresholds online to maximize cache hits while satisfying a user-specified global error-rate constraint. When the system is uncertain, it explores by invoking the underlying model, collects additional correctness observations, and uses them for online threshold learning. This is similar to the approach I would take for Fast answers. In my case, the objective would be to minimize LLM cost while maintaining user satisfaction rather than maximizing cache hits under a global error-rate constraint, and the technical details of learning the policy would also differ.

Selective Question Answering under Domain Shift9 uses gating to answer as many questions as possible while maintaining accuracy above a specified level. Its objective is therefore to maximize coverage subject to a target-accuracy constraint. Because naively using confidence can fail due to calibration problems, the paper trains a calibrator, a type of meta-model, and argues that it remains robust under domain shift. The objective is different, but the approach is similar in that it performs constrained optimization.

Generate-then-Retrieve: Intent-Aware FAQ Retrieval in Product Search10 takes the unusual approach of applying gating before retrieval. This is effective when a large amount of pruning can be done at the user-query level. Although the mechanism is different, exact matching in my design plays a similarly aggressive filtering role: only a narrow subset of queries is allowed to enter the Fast-answer path.

There are many possible decision policies. However, this layer can be made arbitrarily sophisticated, so the design here will keep only the essential pieces and focus on making the system extensible later. Whether a particular method matters should be tested and decided based on the product context. The more important point is to recognize the need for a decision layer and introduce the concept early in the product launch.

4. From Observation to System Design

The rest of this series uses three working assumptions based on the product observations and related work discussed so far.

First, Fast answers are served from a catalog of pre-generated question-answer pairs rather than generated from scratch for every request. Second, to keep the initial system simple and minimize incorrect matches, I will assume that retrieval is based primarily on exact query matching. Third, retrieving a valid QA pair does not necessarily mean that it should be shown. A separate decision layer determines whether serving the Fast answer is preferable to normal LLM generation for a given user and context.

These assumptions are not claims about how ChatGPT is actually implemented. They are design choices I would make if I were responsible for launching the product. With this foundation in place, the next post moves from observation to system design.


If you have read this far, something in this post probably caught your interest. If you would like to know more about me, feel free to connect with me on LinkedIn.

References

  1. See the April 22, 2026 entry in the ChatGPT release notes. ↩ ↩2

  2. Repeated attempts did not trigger a Fast answer for either example. The football question requires fresh information, so it makes sense not to serve a ready-made answer. But why did the Seven Wonders question not trigger the feature? The answer was generated differently each time, while the same four images always appeared. This raises another question: does ChatGPT have a separate caching system for images? A later post will explore that possibility.

    First generated response for Show me the Seven Wonders of the World Second generated response with the same images

    ↩

  3. Korean has a relatively small global speaker base. However, supporting it may be a natural choice given how active South Korea is as a ChatGPT market: in June 2025, The Korea Herald reported that it ranked second globally in paid ChatGPT subscribers, behind only the United States. ↩

  4. See How People Use ChatGPT, Figure 7 and Section 5.2. Seeking Information grew from 14% of messages in July 2024 to 24% in July 2025. ↩

  5. Related work includes AI Knowledge Assist: An Automated Approach for the Creation of Knowledge Bases for Conversational AI Agents (2025); Generating Frequently Asked Questions from Technical Support Tickets using Large Language Models (2025); LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations (2025); DialogQAE: N-to-N Question Answer Pair Extraction from Customer Service Chatlog (2023); and Improving Knowledge Production Efficiency With Question Answering on Conversation (2023). ↩

  6. See How Hyperconnect Built an LLM Explanation Policy. The post is written in Korean, but I recommend reading it with translation. ↩

  7. Related work includes Generate-then-Retrieve: Intent-Aware FAQ Retrieval in Product Search (2023); QUADRo: Dataset and Models for Question-Answer Database Retrieval (2023); Unsupervised FAQ Retrieval with Question Generation and BERT (2020); and Pre-Training Methods for Question Reranking (2024). ↩

  8. See vCache: Verified Semantic Prompt Caching. ↩

  9. See Selective Question Answering under Domain Shift. ↩

  10. See Generate-then-Retrieve: Intent-Aware FAQ Retrieval in Product Search. ↩