SyloSpace

How does ChatGPT work?

ChatGPT is a chatbot built on a large language model that predicts text and is fine-tuned with human feedback, though biased or inaccurate training data can make it unreliable.

Updated 2 hours ago4 min readVersion 2
CommentsFollow

Covers: The basics of how large language models like ChatGPT generate text: training on large text corpora, the transformer architecture, next-token prediction, and fine-tuning with human feedback. Does not cover specific unreleased model internals, pricing, or step-by-step usage guides.

Also answers: How does ChatGPT generate responses? · How do large language models like ChatGPT work? · How is ChatGPT trained?

The short answer

Evidence-backed AI-prepared starting map

ChatGPT is a chatbot built on a large language model (LLM): a neural network trained on a vast amount of text for language tasks, especially generating text. LLMs are typically based on the transformer architecture, and the GPT family is a type of LLM pre-trained to predict the next word, then often fine-tuned to follow instructions and act as an assistant. Beyond pre-training, a software harness can give an LLM memory and access to tools such as web search. Because biased or inaccurate training data can make output less reliable, benchmark evaluations attempt to measure reasoning, factual accuracy, alignment, and safety.1

What this rests on4 independent sources
  • Evidence 16

Did this answer your question?

Be the first to vote

In brief

  1. ChatGPT is a chatbot built on a large language model, typically a transformer neural network trained on a vast amount of text.1

    Evidence-backed
  2. GPT models are pre-trained to predict the next word, then often fine-tuned to follow instructions and behave as assistants.1

    Evidence-backed
  3. Human feedback can be used as a training signal: fine-grained feedback (marking false sentences or irrelevant parts) improved performance on detoxification and long-form question answering.2

    Evidence-backed
  4. Learning from language feedback (ILF) scaled well with dataset size and, combined with comparison feedback, reached human-level summarization performance in one study.3

    Evidence-backed
  5. Biased or inaccurate training data can reduce reliability, which is why benchmarks try to measure reasoning, factual accuracy, alignment, and safety.1

    Evidence-backed

At a glance

What this page stands on

Live · updated just now

The evidence behind it

4 sources
  • Other studies and data3
  • Background1

Published in 2023 and 2024

Sources on this page by kind and year
SourceKindYear
Large language model (Wikipedia)BackgroundUnknown
ChatGPT Alternative Solutions: Large Language Models SurveyOther studies and data2024
Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingOther studies and data2023
Training Language Models with Language Feedback at ScaleOther studies and data2023

The community around it

No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you want the shortest accurate mental model of how ChatGPT produces text

think of it as a transformer-based LLM pre-trained to predict the next word, then fine-tuned to follow instructions and act as an assistant.1

Evidence-backed

If you are trying to understand why assistants behave better than raw next-word predictors

look at fine-tuning with human feedback, where feedback marks specific problems (false sentences, irrelevant parts) and shapes the model's behavior.2

Evidence-backed

If you care about how feedback is collected and used

note that language feedback (ILF) and comparison feedback can be combined, and that combination outperformed either alone in one summarization study.3

Evidence-backed

If you are worried about wrong or biased answers

keep in mind that biased or inaccurate training data can make output less reliable, and that benchmarks attempt to measure factual accuracy, alignment, and safety.1

Evidence-backed

If you want to know what a chatbot can do beyond the base model

consider that a software harness can give an LLM memory and access to tools such as web search.1

Evidence-backed

The full story · 3 chapters

01

What ChatGPT is built on

AI summary:A large language model is a neural network trained on vast text, typically using the transformer architecture, and GPTs are pre-trained to predict the next word then fine-tuned as assistants.

Evidence-backed

Evidence-backed: A large language model is an AI model, typically a neural network, trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can generate, summarize, translate, and analyze text across many contexts, and they underpin many modern chatbots including ChatGPT, Claude, Gemini, Grok, Qwen and DeepSeek. LLMs are typically based on the transformer neural network architecture. Generative pre-trained transformers (GPTs) are a type of LLM pre-trained to predict the next word, and GPTs are then often fine-tuned to follow instructions and behave as assistants.1

Evidence-backed

Evidence-backed: A survey of the LLM field describes ChatGPT as a powerful AI chatbot grounded in LLMs that has drawn widespread societal attention, and frames the technology as reshaping the AI community and the way AI algorithms are created and used. The survey reviews background, key discoveries and prevailing methodologies across multiple LLM models, and identifies existing challenges and potential future research directions.4

02

Training and learning from human feedback

AI summary:Pre-training alone does not make a helpful assistant, so research uses fine-grained human feedback and language feedback as training signals to improve and customize model behavior.

Evidence-backed

Evidence-backed: Pre-training on text teaches a model to predict the next word, but that alone does not make a helpful assistant. One line of research uses fine-grained human feedback as an explicit training signal: instead of a single holistic rating, feedback can mark which sentence is false or which sub-sentence is irrelevant. The Fine-Grained RLHF framework provides a reward after every segment (for example, a sentence) is generated and combines multiple reward models tied to different feedback types such as factual incorrectness, irrelevance, and information incompleteness. Experiments on detoxification and long-form question answering showed improved performance on both automatic and human evaluation, and model behavior could be customized using different combinations of fine-grained reward models.2

Evidence-backed

Evidence-backed: A second approach, Imitation learning from Language Feedback (ILF), uses richer language feedback rather than only comparisons. It iterates three steps: condition the model on the input, an initial output, and feedback to generate refinements; select the refinement that incorporates the most feedback; then fine-tune the model to maximize the likelihood of the chosen refinement given the input. The authors show theoretically that ILF can be viewed as Bayesian inference, similar to reinforcement learning from human feedback. On a controlled toy task and a realistic summarization task, large language models accurately incorporated feedback, and fine-tuning with ILF scaled well with dataset size, even outperforming fine-tuning on human summaries. Learning from both language and comparison feedback outperformed learning from each alone, reaching human-level summarization performance.3

03

Limits and how models are evaluated

AI summary:Biased or inaccurate training data can reduce reliability, so benchmarks measure reasoning, factual accuracy, alignment, and safety, while a software harness can add memory and tools like web search.

Evidence-backed

Evidence-backed: Biased or inaccurate training data can make an LLM's output less reliable. To gauge quality, benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety. A software harness can also provide LLMs with memory and access to tools such as web search, extending what the base model can do on its own.1

Your turn

Have your say

Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.

How do you feel about this?

No votes yet

Quick questions from connected pages

Before you go

What to remember

The few things worth keeping from this page.

  1. ChatGPT is a chatbot built on a large language model, typically a transformer neural network trained on a vast amount of text.

  2. GPT models are pre-trained to predict the next word, then often fine-tuned to follow instructions and behave as assistants.

  3. Human feedback can be used as a training signal: fine-grained feedback (marking false sentences or irrelevant parts) improved performance on detoxification and long-form question answering.

This answer keeps changing

When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 2 hours ago). Follow it to be told when that happens.

Up nextHow does OpenAI watermark ChatGPT outputs in the EU?How does OpenAI watermark ChatGPT outputs in the EU, and what does the EU AI Act require?

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    Large language model (Wikipedia)
    WikipediaPublished Oct 10, 2026Checked Oct 10, 2026
    “A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts. They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, Qwen and DeepSeek. LLMs are typically based on the transformer neural network architecture. Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word. GPTs are then often fine-tuned to follow instructions and to behave as assistants. Biased or inaccurate training data can make an LLM's output less reliable. A software harness can provide LLMs with memory and access to tools such as web search. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety.”
  2. 2
    Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
    arXiv (Cornell University) (Wu et al.)Published Jun 2, 2023Checked Oct 10, 2026
    “However, such holistic feedback conveys limited information on long text outputs; it does not indicate which aspects of the outputs influenced user preference; e.g., which parts contain what type(s) of errors. In this paper, we use fine-grained human feedback (e.g., which sentence is false, which sub-sentence is irrelevant) as an explicit training signal. We introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.g., a sentence) is generated; and (2) incorporating multiple reward models associated with different feedback types (e.g., factual incorrectness, irrelevance, and information incompleteness). We conduct experiments on detoxification and long-form question answering to illustrate how learning with such reward functions leads to improved performance, supported by both automatic and human evaluation. Additionally, we show that LM behaviors can be customized using different combinations of fine-grained reward models. We release all data, collected human feedback, and codes at https://FineGrainedRLHF.github.io.”
  3. 3
    Training Language Models with Language Feedback at Scale
    arXiv (Cornell University) (Scheurer et al.)Published Mar 28, 2023Checked Oct 10, 2026
    “However, comparison feedback only conveys limited information about human preferences. In this paper, we introduce Imitation learning from Language Feedback (ILF), a new approach that utilizes more informative language feedback. ILF consists of three steps that are applied iteratively: first, conditioning the language model on the input, an initial LM output, and feedback to generate refinements. Second, selecting the refinement incorporating the most feedback. Third, finetuning the language model to maximize the likelihood of the chosen refinement given the input. We show theoretically that ILF can be viewed as Bayesian Inference, similar to Reinforcement Learning from human feedback. We evaluate ILF's effectiveness on a carefully-controlled toy task and a realistic summarization task. Our experiments demonstrate that large language models accurately incorporate feedback and that finetuning with ILF scales well with the dataset size, even outperforming finetuning on human summaries. Learning from both language and comparison feedback outperforms learning from each alone, achieving human-level summarization performance.”
  4. 4
    ChatGPT Alternative Solutions: Large Language Models Survey
    Research paper (Alipour et al.)Published Mar 16, 2024Checked Oct 10, 2026
    “Recent years have witnessed a dynamic synergy between academia and industry, propelling the field of LLM research to new heights. A notable milestone in this journey is the introduction of ChatGPT, a powerful AI chatbot grounded in LLMs, which has garnered widespread societal attention. The evolving technology of LLMs has begun to reshape the landscape of the entire AI community, promising a revolutionary shift in the way we create and employ AI algorithms. Given this swift-paced technical evolution, our survey embarks on a journey to encapsulate the recent strides made in the world of LLMs. Through an exploration of the background, key discoveries, and prevailing methodologies, we offer an up-to-theminute review of the literature. By examining multiple LLM models, our paper not only presents a comprehensive overview but also charts a course that identifies existing challenges and points toward potential future research trajectories. This survey furnishes a well-rounded perspective on the current state of generative AI, shedding light on opportunities for further exploration, enhancement, and innovation.”

How it changed

Published 1 time since Oct 10, 2026.

  1. Version 2Oct 10, 2026Live now

    AI-prepared Starting Map from live research.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

  • “Limits and how models are evaluated” rests on one independent source

    A second, independent source that confirms or challenges it would make this part more reliable.

Open questions

  • How closely do the published fine-tuning methods (fine-grained RLHF, ILF) match the training process actually used for ChatGPT?

    No answers yet

  • What text corpora and scale of compute were used to pre-train ChatGPT, and how do those choices affect its outputs?

    No answers yet

  • How well do current benchmarks capture real-world reliability, bias, and safety of assistant-style models?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Ask this Sylo

Answers only from “How does ChatGPT work?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.