Enter a question. Scroll down to understand every physical layer that has to work before AI can answer.
Your question is broken into tokens, encrypted, and sent as a tiny data packet toward a data center that might be thousands of miles away.
Everything you type becomes numbers before it leaves your hand
Your prompt races through glass fiber at two-thirds the speed of light, bouncing off satellites or snaking along the ocean floor through 800,000+ miles of submarine cable.
Your question crosses an ocean faster than you can blink
Your prompt is milliseconds away from a data center — but that data center only exists because someone found it a gigawatt of electricity. Grid connections now take longer to get than the chips do.
The constraint on AI is no longer chips — it's electricity
Your prompt arrives at a building the size of several football fields. Big Tech is spending over $600 billion in 2026 building more of them.
AI lives in buildings that draw as much power as cities
Your prompt gets split across 72 GPUs that talk to each other at terabytes per second.
72 chips pretend to be one giant brain
The rack isn't running a mind — it's running a loop. Read the whole question at once, then predict one token at a time, each one pulled through memory at terabytes per second.
This is why AI answers appear word by word
Each GPU has 336 billion transistors flanked by towers of stacked memory feeding data at up to 22 terabytes per second.
Your answer is limited by how fast memory can feed the math
TSMC makes virtually all AI chips. A single dust particle can ruin a chip with billions of transistors smaller than a virus.
One chip takes three months and thousands of steps to make
This is the foundation of everything. A laser hits tin, mirrors shape the light, and the pattern is printed onto a wafer. Only ~70 of these machines are built per year.
Every AI answer traces back to ~70 machines a year from one company
EUV light prints 70+ circuit layers, each aligned to sub-nanometer precision — a few atoms wide.
The wafer is cut, tested, and paired with HBM memory stacks on a single substrate.
Terabytes per second of bandwidth stream your context through stacked DRAM towers into the GPU die.
One token at a time, the loop runs: predict, append, repeat — until your answer exists.
72 coordinated GPUs turn your question into the math that produces each next word.
Every token generated leans on a gigawatt-scale power draw that the grid quietly sustains.
The generated tokens leave the data center as packets, headed back across the network.
At two-thirds the speed of light, back across the ocean and into your local network.
Your device decrypts the tokens and the answer appears, character by character.
That answer depended on fiber networks, data centers, GPUs, advanced memory, and chip fabs running at extreme precision.
Energy and cost are rough estimates. Reasoning models can use 10–100× more.