SyloSpace

How much energy does ChatGPT use?

A generative AI reply can use far more energy than a non-generative path, and prompt and infrastructure choices change the footprint a lot.

Updated 53 minutes ago5 min readVersion 2
CommentsFollow

Covers: Published estimates of the electricity and water use of ChatGPT and similar large language models, including per-query figures, training costs, and how those estimates are calculated. It does not cover the energy use of other AI applications in detail or provide a full life-cycle assessment of data centres.

Also answers: How much electricity does ChatGPT use per prompt? · What is ChatGPT's carbon footprint? · How much water does ChatGPT use? · ChatGPT energy consumption compared to Google search

Rows of black server racks with white logos in a data center
Photo: imgix

Before you read, make a guess

Fill in the blank: ?% of inference energy and emissions cut by best-practice prompt engineering

Drag the slider to fill in the blank

0%50%100%

The short answer

Evidence-backed AI-prepared starting map

Published estimates put the energy of a single generative AI reply in the range of roughly 168 mWh (0.168 Wh) in one clinical evaluation, while a non-generative path on the same local hardware used about 2.23 mWh per request — about 75 times less. Estimates vary widely with model, prompt and hardware, and prompt design alone can shift inference energy and emissions by 32–48%. At population scale, extrapolating general-practice survey data suggests ChatGPT queries in NHS primary care alone could release about 349 t CO2e per year.123

What this rests on5 independent sources
  • Evidence 15
  • Interpretation 1

Did this answer your question?

Be the first to vote

In brief

  1. One clinical evaluation measured about 168.27 mWh per generative reply versus about 2.23 mWh for a non-generative path on the same hardware — roughly 75 times more energy for the generative reply.1

    Evidence-backed
  2. Better prompt design cut inference energy and operational emissions by 32–48% across a range of LLMs and scenarios.2

    Evidence-backed
  3. Extrapolating survey data, ChatGPT queries in NHS primary care alone could release about 349 t CO2e per year.3

    Evidence-backed
  4. Infrastructure choices matter: renewable-backed, lower-PUE infrastructure cut modelled operational emissions by about 79% in one large-scale scenario.4

    Evidence-backed
  5. Per-query figures are not one number: they depend on the model, the prompt, the hardware and whether the task needs generation at all.12

    Interpretation

At a glance

The picture in numbers

Live · updated just now

Same local hardware in one clinical evaluation
  • Generative reply168.27 mWh
  • Non-generative path2.23 mWh

Generative reply is about 75 times non-generative path.

Energy per request: generative reply vs non-generative path1
Across a range of LLMs and test scenarios

32%

32 in every 100

of inference energy and emissions cut by best-practice prompt engineering2
Extrapolated from general-practice survey data

349 t CO2e per year

CO2e per year from ChatGPT queries in NHS primary care3
One large-scale scenario

79%

79 in every 100

of modelled operational emissions cut by renewable-backed, lower-PUE infrastructure4

The evidence behind it

5 sources
  • Other studies and data4
  • Background1

Published in 2025 and 2026

Sources on this page by kind and year
SourceKindYear
A Locally Executable AI System for Improving Preoperative Patient Communication: Multidomain Clinical Evaluation.Other studies and data2026
Carbon Reporting Practices in the NHS: Emissions and Omissions Relating to Artificial Intelligence.Other studies and data2025
Green prompt engineering for sustainable generative AI.Other studies and data2026
Quantifying and Disclosing the Environmental Footprint of AI in Research: Life Cycle-Informed Framework and Open-Access Calculator Development Study.Other studies and data2026
Large language model (Wikipedia)BackgroundUnknown

The community around it

No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.

Shares and multiples are worked out from the figures the page states.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you are estimating your organisation's AI footprint

count unprocured generative AI tools as well as owned hardware, and push for AI-specific carbon disclosure clauses in vendor contracts and cradle-to-grave hardware factors in Scope 3 reporting, since these are the gaps identified in NHS carbon reporting.3

Evidence-backed

If you are designing a generative AI application

apply best-practice prompt engineering to avoid repeated iterations; across a range of LLMs and scenarios this cut energy use and operational emissions by 32–48%.2

Evidence-backed

If your task does not need free-form generation

a non-generative retrieval path can use far less energy — about 2.23 mWh per request versus about 168.27 mWh for the generative path in one clinical evaluation, roughly 75 times less.1

Evidence-backed

If you are choosing where AI workloads run

renewable-backed, lower-PUE infrastructure cut modelled operational emissions by about 79% in a large-scale scenario, though the same analysis notes a water–carbon trade-off that should be reported.4

Evidence-backed

If you need a defensible number for a report

use an auditable calculator with explicit energy, emissions, water, hardware and regional assumptions, as routine disclosure of these parameters is what makes AI footprint figures reproducible.4

Evidence-backed

The full story · 3 chapters

01

What a single query is estimated to use

Evidence-backed

Evidence-backed: In a multidomain clinical evaluation, the generative path (a small language model handling 'small talk') used approximately 168.27 mWh per request, with latency of about 8.51 seconds and video RAM averaging about 13.3 GiB (peak about 14.0 GiB). The non-generative clinical path on the same local, low-cost hardware used approximately 2.23 mWh per request, with latency about 0.10 seconds and video RAM averaging about 2.2 GiB (peak about 2.5 GiB) — roughly a 75-fold higher energy footprint per reply for the generative path. The authors present decoupling clinical information retrieval from generative chitchat as a way to cut energy use while preserving safety and privacy.1

Evidence-backed

Evidence-backed: Prompt design is one of the levers on per-query energy. Across a range of LLMs and test scenarios, applying best-practice prompt engineering reduced energy consumption and the corresponding operational greenhouse gas emissions by 32–48%, mainly by avoiding the repeated iterations that suboptimal prompts require. The authors propose integrating these practices into the design of generative AI applications.2

Readers' pollNo answers yet

How concerned are you about the energy use of AI tools like ChatGPT?

How concerned are you about the energy use of AI tools like ChatGPT?

Your individual answer is private. Only totals are shown.

02

From one query to a whole organisation

AI summary:Scaling per-query figures to organisations depends on assumptions, and optimisation and cleaner infrastructure can cut emissions substantially.

Evidence-backed

Evidence-backed: Scaling per-query figures up is where the numbers get large and the assumptions get load-bearing. Extrapolating general practice survey data, ChatGPT queries alone could release approximately 349 t CO2e per year in NHS primary care. The same analysis warns that carbon reporting in the NHS lacks granularity, so averages can obscure the extreme energy intensity of certain AI workloads; that life-cycle emissions from specialised hardware such as GPUs are often excluded unless the trust owns the equipment, ignoring upstream manufacturing impacts; and that widespread use of unprocured generative AI tools is unmeasured. Proposed fixes include AI-specific carbon disclosure clauses in vendor contracts, cradle-to-grave emission factors for AI hardware in Scope 3 reporting, and lightweight monitoring of external AI traffic (with acknowledged ethical concerns).3

Evidence-backed

Evidence-backed: A life-cycle-informed framework and open-access calculator offers a way to make these estimates auditable. In its scenarios, optimisation cut total emissions from about 0.07 kg CO2e to 0.05 kg CO2e in a small laboratory (improving the proposed label from C to B); carbon-aware scheduling cut a midsized cluster from about 665 kg CO2e/month to 450 kg CO2e/month, a 32% reduction; and shifting to renewable-backed, lower-PUE infrastructure cut operational emissions by about 79% in the large-scale scenario while increasing the importance of water–carbon trade-off reporting. The authors argue routine disclosure of energy use, emissions, water use, hardware assumptions and regional context improves reproducibility.4

03

What ChatGPT is, in energy terms

AI summary:ChatGPT is a large language model, and the same visible query can trigger very different amounts of computation.

Evidence-backed

Evidence-backed: ChatGPT is a chatbot built on a large language model — typically a transformer neural network trained on a vast amount of text, and in the GPT case pre-trained to predict the next word and then fine-tuned to follow instructions. LLMs can generate, summarise, translate and analyse text, and a software harness can give them memory and access to tools such as web search. This matters for energy accounting because the same user-visible 'query' can trigger very different amounts of computation depending on the model, the harness and whether tools are called.5

Your turn

Have your say

Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.

How do you feel about this?

No votes yet

Quick questions from connected pages

Before you go

What to remember

Try to recall each hidden figure before you reveal it. Remembering, not rereading, is what makes it stick.

  1. One clinical evaluation measured about mWh per generative reply versus about 2.23 mWh for a non-generative path on the same hardware — roughly 75 times more energy for the generative reply.

  2. Better prompt design cut inference energy and operational emissions by across a range of LLMs and scenarios.

  3. Extrapolating survey data, ChatGPT queries in NHS primary care alone could release about t CO2e per year.

This answer keeps changing

When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 53 minutes ago). Follow it to be told when that happens.

Up nextHow much electricity do data centres and AI use?How much electricity do data centres and artificial intelligence use, and how is that demand expected to change?

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    A Locally Executable AI System for Improving Preoperative Patient Communication: Multidomain Clinical Evaluation.
    JMIR medical informatics (Sato et al.)Published Jul 21, 2026Checked Oct 10, 2026
    “This performance was not statistically different from that of ChatGPT (GPT-4o), which had 6 out of 400 (1.5%) errors (McNemar test with Holm adjustment; P>.99). Sustainability measurements showed approximately 2.23 mWh per request (latency≈0.10 s; video RAM≈2.2 GiB average, ≈2.5 GiB peak) for the nongenerative clinical path vs approximately 168.27 mWh (latency≈8.51 s; video RAM≈13.3 GiB average, ≈14.0 GiB peak) for small language model small talk-approximately a 75-fold higher energy footprint per reply for the generative path.ConclusionsHigh-precision, nongenerative clinical support is feasible using local, low-cost hardware without cloud dependence. By decoupling clinical information retrieval from generative chitchat, LENOHA enhances safety, preserves privacy, and markedly reduces energy use, offering a practical blueprint for sustainable and equitable medical AI deployment across diverse care settings.”
  2. 2
    Green prompt engineering for sustainable generative AI.
    Environmental science and ecotechnology (Podder et al.)Published Mar 13, 2026Checked Oct 10, 2026
    “However, substantial computational costs and energy footprint of prompt inferencing process remain critical challenges while building generative AI applications. The energy efficiency of LLM inferences is particularly impacted by suboptimal prompts, which may require multiple iterations, thereby escalating energy consumption and the associated carbon footprint. To address these challenges, we propose a series of practices and guidelines designed to enhance the likelihood of obtaining desired responses from LLMs with minimal reiterations. Empirical evaluation demonstrates that, across a range of LLMs and test scenarios, energy consumption and corresponding operational greenhouse gas emissions were reduced by 32-48% when best practices were applied. Drawing upon these insights, our proposed best practices can be seamlessly integrated into the design frameworks of generative AI applications, thereby enhancing the energy efficiency of prompt inferencing. By addressing the challenge of establishing a cohesive framework for energy-efficient prompt design and inferencing, this paper advocates for the sustainable and effective deployment of generative AI technologies.”
  3. 3
    Carbon Reporting Practices in the NHS: Emissions and Omissions Relating to Artificial Intelligence.
    Journal of medical Internet research (Reynolds)Published Oct 27, 2025Checked Oct 10, 2026
    “First, a lack of granularity provides averages that can obscure the extreme energy intensity of certain AI workloads. Second, life-cycle emissions from specialized hardware (eg, graphics processing units) are often excluded unless trusts own the equipment, ignoring upstream manufacturing impacts. Third, widespread use of unprocured generative AI tools is unmeasured; extrapolating general practice survey data suggests that ChatGPT queries alone could release ≈ 349t CO₂e per year in primary care. To close these gaps, we propose three potential ways to help reduce these reporting gaps: (1) AI-specific carbon disclosure clauses in vendor contracts, (2) inclusion of cradle-to-grave emission factors for AI hardware in Scope 3 reporting, and (3) lightweight monitoring of external AI traffic (while recognizing potential ethical issues with this). Implementing these measures would give health care leaders a more accurate baseline against which to judge whether AI supports or undermines the NHS net-zero target.”
  4. 4
    Quantifying and Disclosing the Environmental Footprint of AI in Research: Life Cycle-Informed Framework and Open-Access Calculator Development Study.
    JMIR AI (Mozafari et al.)Published Aug 18, 2026Checked Oct 10, 2026
    “In the small laboratory scenario, optimization reduced total emissions from approximately 0.07 kg carbon dioxide equivalents (CO2e) to 0.05 kg CO2e and improved the proposed label from C to B. In the midsized cluster scenario, carbon-aware scheduling reduced emissions from approximately 665 kg CO2e/month to 450 kg CO2e/month-a 32% reduction. In the large-scale scenario, shifting to renewable-backed, lower-PUE infrastructure reduced operational emissions by approximately 79% while increasing the importance of water-carbon trade-off reporting.ConclusionsThe calculator provides a practical and transparent method for reporting AI environmental footprints using auditable parameters and publication-ready outputs. Routine disclosure of energy use, emissions, water use, hardware assumptions, and regional context can improve reproducibility and support more equitable and sustainable AI research evaluation.”
  5. 5
    Large language model (Wikipedia)
    WikipediaPublished Oct 10, 2026Checked Oct 10, 2026
    “A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts. They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, Qwen and DeepSeek. LLMs are typically based on the transformer neural network architecture. Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word. GPTs are then often fine-tuned to follow instructions and to behave as assistants. Biased or inaccurate training data can make an LLM's output less reliable. A software harness can provide LLMs with memory and access to tools such as web search. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety.”

How it changed

Published 1 time since Oct 10, 2026.

  1. Version 2Oct 10, 2026Live now

    AI-prepared Starting Map from live research.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

  • “What ChatGPT is, in energy terms” rests on one independent source

    A second, independent source that confirms or challenges it would make this part more reliable.

Open questions

  • How much energy and carbon go into training a model like GPT-4o, and how does that compare with the cumulative energy of its queries?

    No answers yet

  • How does a per-query figure of roughly 168 mWh compare with common everyday activities such as boiling a kettle, a phone charge or a search-engine query?

    No answers yet

  • What is the measured per-query energy of ChatGPT itself, as opposed to a small language model running on local hardware?

    No answers yet

  • How much water does a ChatGPT query consume, and how do water and carbon trade off against each other?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Ask this Sylo

Answers only from “How much energy does ChatGPT use?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.