Scroll to explore

Prompt to Chip

Enter a question. Scroll down to understand every physical layer that has to work before AI can answer.

2026
[ The Application Layer ]
Your prompt leaves the device as a data packet
Your question here YOUR DEVICE TOKENIZES, ENCRYPTS, SENDS DATA PACKET HEADING TO THE NETWORK OPENAI · ANTHROPIC · GOOGLE · META

Your question is broken into tokens, encrypted, and sent as a tiny data packet toward a data center that might be thousands of miles away.

Everything you type becomes numbers before it leaves your hand

ENCRYPTED PACKET LEAVES DEVICE
WiFiWiFi The wireless link between your device and a nearby router, typically the first hop your data packet takes.
ENTERS THE NETWORK
2026
[ The Network Layer ]
Your prompt travels through fiber, ground stations, and satellites
FIBER OPTIC CABLE GLASS CORE — YOUR DATA RIDES THIS CLADDING — KEEPS LIGHT IN PROTECTIVE JACKET SUBMARINE CABLE 800,000+ MILES ON THE OCEAN FLOOR LEO SATELLITE ALTERNATE ROUTE TRANSMISSION SPEED FIBER: ~200,000 km/s LATENCY: 10-50 ms CLOUDFLARE · EQUINIX · SPACEX · AT&T

Your prompt races through glass fiber at two-thirds the speed of light, bouncing off satellites or snaking along the ocean floor through 800,000+ miles of submarine cable.

Your question crosses an ocean faster than you can blink

FIBER TRUNK
GRIDThe Grid Before your prompt can be computed, the data center it's heading to has to be powered — AI now strains regional electric grids.
THE POWER LAYER
2026
[ The Power Layer ]
The new bottleneck — five AI campuses cross 1 gigawatt in 2026
YOUR PROMPT ARRIVES GENERATION GAS TURBINES · NUCLEAR · SOLAR TRANSMISSION · 115kV+ SUBSTATION + TRANSFORMERS 115kV → 13.8kV → 415V SCARCE · MULTI-YEAR LEAD TIMES 1 GW CAMPUS ≈ ONE NUCLEAR REACTOR POWER FLOW → THE QUEUE GRID CONNECTION WAIT: 4–7 YEARS ~HALF OF PLANNED US SITES DELAYED 1 GW CLUB (2026) STARGATE ABILENE · COLOSSUS 2 PROMETHEUS · FAYETTEVILLE · NEW CARLISLE POWER SECURED · THE PROMPT CAN NOW BE COMPUTED GE VERNOVA · SIEMENS ENERGY · CONSTELLATION · TVA

Your prompt is milliseconds away from a data center — but that data center only exists because someone found it a gigawatt of electricity. Grid connections now take longer to get than the chips do.

The constraint on AI is no longer chips — it's electricity

SUBSTATION → RACK
115kV115 kilovolts High-voltage transmission feeds a dedicated on-site substation, which steps power down to what servers can use.
ARRIVES AT DATA CENTER
2026
[ The Data Center ]
A warehouse burning megawatts of power — your prompt arrives here
YOUR PROMPT ARRIVES SERVER RACKS 100,000+ GPUs AT FRONTIER SITES COOLING EVAPORATIVE TOWERS REJECT FACILITY HEAT COOLING EVAPORATIVE TOWERS REJECT FACILITY HEAT POWER SUBSTATION 100 MW – 1 GW+ FIVE 1 GW SITES IN 2026 GRID → RACK LIQUID COOLING CLOSED-LOOP COOLANT DIRECT-TO-CHIP COOLANT AIR CAN'T KEEP UP SCALE (2026) $600B+ BIG TECH CAPEX ≈ $1.7B EVERY DAY AWS · GOOGLE CLOUD · MICROSOFT AZURE · COREWEAVE

Your prompt arrives at a building the size of several football fields. Big Tech is spending over $600 billion in 2026 building more of them.

AI lives in buildings that draw as much power as cities

ROUTED TO GPU CLUSTER
IB 400GInfiniBand 400G An ultra-fast interconnect used inside data centers so GPUs can exchange data with very low latency.
INFINIBAND
2026
[ The GPU Rack — NVL72 ]
72 GPUs working as one brain, connected at terabytes per second
GPU MODULE 8x GPU DIES + HBM NVLink BACKBONE 1.8 TB/s BETWEEN GPUs NVL72 RACK 72 GPUs PER RACK ~120 kW POWER ~$3M PER RACK NEXT: RUBIN NVL144 · H2 2026 NVIDIA · DELL · SUPER MICRO · VERTIV

Your prompt gets split across 72 GPUs that talk to each other at terabytes per second.

72 chips pretend to be one giant brain

72 GPUs · ONE MODEL
MATHThe Computation Everything below this line is physics and matrix multiplication — the model itself is just numbers arranged by training.
WHAT THE GPUs COMPUTE
2026
[ Inside the Model ]
What those 72 GPUs actually compute — and why answers stream word by word
1 · YOUR WORDS BECOME TOKENS Why is the sky blue ? 2 · PREFILL — READ EVERYTHING AT ONCE ALL TOKENS AT ONCE · ONE PASS 3 · DECODE — ONE TOKEN AT A TIME REPEAT FOR EVERY WORD PREDICT NEXT TOKEN NEW TOKEN REASONING MODELS (2026) MAY THINK IN 1,000s OF HIDDEN TOKENS BEFORE THE FIRST VISIBLE WORD KV CACHE REMEMBERS WHAT'S ALREADY READ SO EACH WORD ISN'T RECOMPUTED GPT · CLAUDE · GEMINI · LLAMA

The rack isn't running a mind — it's running a loop. Read the whole question at once, then predict one token at a time, each one pulled through memory at terabytes per second.

This is why AI answers appear word by word

NVLink / PCIe
PCIePCIe The high-speed bus inside a server that connects the CPU, GPUs, and other accelerators.
INTO THE SILICON
2026
[ The GPU Die + High Bandwidth Memory ]
Inside NVIDIA's newest Rubin-generation GPU — where inference happens
CoWoS PACKAGE SUBSTRATE HBM4 HBM4 HBM4 HBM4 GPU COMPUTE DIE 336B TRANSISTORS · TSMC 3nm HBM4 STACK 8 STACKS · 288 GB · 22 TB/s NVIDIA · AMD · GOOGLE TPU · SK HYNIX · MICRON

Each GPU has 336 billion transistors flanked by towers of stacked memory feeding data at up to 22 terabytes per second.

Your answer is limited by how fast memory can feed the math

WHERE THESE CHIPS COME FROM
SiSilicon (Si) The base semiconductor material that nearly all modern chips are built on.
THE SEMICONDUCTOR FAB
2026
[ The Semiconductor Fab ]
The most complex buildings humans construct
FAN FILTER UNITS · AIR REPLACED EVERY 3 SECONDS PROCESS TOOL ION IMPLANT / CVD MODIFIES SILICON SURFACE ETCH TOOL PLASMA CARVES CIRCUITS NANOMETER ACCURACY DEPOSITION CVD / ALD THIN FILMS BUILDS LAYERS ATOM BY ATOM METROLOGY INSPECTS EVERY LAYER DEFECT = SCRAPPED WAFER FOUP FOUP FAB CLEANROOM 3-5 YEARS TO BUILD · $20-50B CAPEX TSMC · SAMSUNG FOUNDRY · INTEL FOUNDRY

TSMC makes virtually all AI chips. A single dust particle can ruin a chip with billions of transistors smaller than a virus.

One chip takes three months and thousands of steps to make

THE MACHINE INSIDE THE FAB
EUVExtreme Ultraviolet Light with a 13.5nm wavelength used to print the tiniest transistor features on advanced chips. Only ASML makes these machines.
13.5nm LIGHT
2026
[ ASML EUV Lithography ]
The deepest layer. $200–400M per machine. ~70 built per year.
VACUUM CHAMBER LASER + TIN SOURCE CO2 PULSE HITS TIN PLASMA CREATES EUV LIGHT 50,000 PULSES/SEC MIRROR TRAIN ZEISS OPTICS SHAPE THE BEAM POLISHED TO <0.05nm SMOOTHNESS M1 M2 M3 MASK + WAFER STAGE nm-PRECISION POSITIONING THE BOTTLENECK ~70 TOOLS / YEAR $200-400M EACH SOLE MAKER: ASML EUV LIMITS CHIP SUPPLY NEXT: HIGH-NA EUV ~$400M ASML · ZEISS · CYMER THE MACHINE THAT SETS THE CEILING FOR AI

This is the foundation of everything. A laser hits tin, mirrors shape the light, and the pattern is printed onto a wafer. Only ~70 of these machines are built per year.

Every AI answer traces back to ~70 machines a year from one company

Now the response races back up through every layer
09

Wafer Patterned

EUV light prints 70+ circuit layers, each aligned to sub-nanometer precision — a few atoms wide.

08

Chip Packaged

The wafer is cut, tested, and paired with HBM memory stacks on a single substrate.

07

Memory Feeds the Model

Terabytes per second of bandwidth stream your context through stacked DRAM towers into the GPU die.

06

Model Predicts

One token at a time, the loop runs: predict, append, repeat — until your answer exists.

05

GPUs Generate Tokens

72 coordinated GPUs turn your question into the math that produces each next word.

04

The Grid Holds

Every token generated leans on a gigawatt-scale power draw that the grid quietly sustains.

03

Data Center Sends

The generated tokens leave the data center as packets, headed back across the network.

02

Fiber Returns

At two-thirds the speed of light, back across the ocean and into your local network.

01

Rendered on Screen

Your device decrypts the tokens and the answer appears, character by character.

What is the meaning of life?
AI

That answer depended on fiber networks, data centers, GPUs, advanced memory, and chip fabs running at extreme precision.

Your Question’s Footprint

Energy and cost are rough estimates. Reasoning models can use 10–100× more.

Words you now know
Token Latency Submarine Cable Gigawatt Interconnection Queue NVLink Prefill Decode KV Cache HBM CoWoS EUV