AI

LLMs as Probability Machines — The Elegance of Next Token Prediction

TL;DR Large Language Models are complex mathematical functions that predict the next probable token in a sequence. But that simple description masks breathtaking depth: hundreds of billions of learned parameters, parallel transformer attention mechanisms, reinforcement learning from human feedback, and conversational capabilities that emerge from the architecture in ways nobody fully anticipated. This is the complete, honest engineering explanation.

Read this article as text (accessible version)
AI Fundamentals — Deep Dive

How Large Language Models Actually Work: The Engineering Truth Behind the Magic

📅 June 2025⏱ ~21 min read🎯 3,800+ words

You've used ChatGPT, Gemini, or Claude. You've watched them write poetry, debug code, and explain quantum physics. But do you actually know what's happening inside? Not the marketing version — the real engineering. This guide pulls the curtain all the way back: probability distributions, transformer attention, backpropagation, and the emergent behavior that even the researchers who built these systems can't fully explain.

Next Token Prediction Parameters & Training RLHF Transformer Attention Emergent Behavior Post Excerpt

Large Language Models are complex mathematical functions that predict the next probable token in a sequence. But that simple description masks breathtaking depth: hundreds of billions of learned parameters, parallel transformer attention mechanisms, reinforcement learning from human feedback, and conversational capabilities that emerge from the architecture in ways nobody fully anticipated. This is the complete, honest engineering explanation.

Introduction

The Autocomplete That Fooled Everyone

In November 2022, a software engineer at a mid-sized tech company was writing a unit test when he decided to try the new ChatGPT thing. He typed half a function, hit enter, and watched it complete not just the code but the entire test suite — including edge cases he hadn't thought of. He sat back, slightly unsettled, and asked it to explain the algorithm it had just written. It explained it flawlessly. He asked a follow-up. It answered that too. He thought: "This is not what I expected autocomplete to be."

That engineer was experiencing something that millions of people would experience over the following months: the gap between expectation and reality when encountering a large language model for the first time. The expectation was clever autocomplete. The reality felt different — more like a conversation with something that understood. But understanding the gap between that feeling and the actual engineering underneath it is one of the most important things you can do as a developer, researcher, or product builder in the current era of AI. Because the engineering reality is both more impressive and more humbling than the surface experience suggests.

The Core Misconception

LLMs are not "thinking" in any sense that resembles human cognition. They don't reason, plan, or understand in the way those words imply. They are extraordinarily sophisticated statistical prediction engines — and that description is not a dismissal. Producing the outputs these models produce purely through mathematical prediction of probable next tokens, without any symbolic reasoning or explicit knowledge base, is an astonishing engineering achievement. The magic isn't compromised by understanding the mechanism. It's deepened by it.

Core Mechanism

LLMs as Probability Machines — The Elegance of Next Token Prediction

Strip away every press release and every breathless headline, and you'll find at the core of every large language model one simple, powerful task: given a sequence of tokens (words, subwords, punctuation marks), predict the probability distribution over every possible next token. That's it. The model doesn't maintain a running plan. It doesn't have a goal beyond filling the next slot. It takes in a sequence, consults its hundreds of billions of learned parameters, and outputs a probability for every token in its vocabulary (typically 50,000–100,000 tokens).

The generation process is iterated. You sample a token from that distribution (or take the highest-probability one, depending on temperature settings), append it to the sequence, feed the new longer sequence back through the model, and get probabilities for the next token. Token by token, a coherent response assembles. The remarkable thing isn't that this works — the remarkable thing is how well it works. The statistical regularities in human language are so rich that a model trained purely to predict next tokens ends up learning grammar, world knowledge, logical reasoning patterns, writing style, and domain expertise as instrumental side effects of doing its core job well.

Tokens aren't exactly words — they're subword units produced by algorithms like Byte Pair Encoding (BPE). "Unbelievable" might tokenize as ["Un", "believe", "able"]. Common words like "the" are single tokens; rare words might be split into many. This matters because the context window limit in tokens translates to fewer actual words than you might expect, and token count drives API costs directly. Understanding tokenization is the first step to understanding why LLMs sometimes behave unexpectedly on rare words or very long documents.

Real-World Analogy

Imagine a musician who has spent decades listening to every piece of music ever recorded. They can't describe music theory explicitly — they've never studied it. But if you play the first eight bars of a Beethoven sonata and ask them to hum the next bar, they'll produce something remarkably Beethoven-like. They've internalized the statistical patterns of music so deeply that their prediction of "what comes next" is indistinguishable from composition. That musician is an LLM. The corpus of human text is their lifetime of listening. Their hummed next bar is the predicted next token.

LLM next token prediction probability distribution diagram showing vocabulary tokens with probability bars and sampling mechanism
# Next token prediction with temperature — what actually happens under the hood
import torch
import torch.nn.functional as F

def sample_next_token(logits: torch.Tensor, temperature: float = 1.0) -> int:
 """
 logits: raw output scores for every vocabulary token [vocab_size]
 temperature: 0 = always pick highest prob, 1 = sample from distribution
 """
 # Scale logits by temperature (lower = more peaked distribution)
 scaled = logits / temperature

 # Convert to probabilities via softmax
 probs = F.softmax(scaled, dim=-1)

 # Sample one token from the distribution
 next_token = torch.multinomial(probs, num_samples=1)
 return next_token.item()

# Temperature = 0.1: nearly deterministic (picks highest prob token)
# Temperature = 1.0: sampling from true distribution
# Temperature = 1.5: exaggerated randomness, more "creative" but incoherent
⚡ Pro Tips — Understanding Token Prediction
  • Temperature is the single most important parameter to understand when working with LLMs. Low temperature (0–0.3) gives you deterministic, consistent outputs — good for code generation, data extraction, structured tasks. Higher temperature (0.7–1.0) enables creativity and variety — good for brainstorming, writing, exploration.
  • Token limits are real constraints, not suggestions. A 128k context window sounds enormous but fills up fast with long conversations, system prompts, and documents. Always count tokens explicitly using tiktoken (OpenAI) or equivalent — don't estimate.
  • Here's the thing most tutorials miss: the model doesn't "read" your prompt in order — it processes all tokens in parallel through the attention mechanism, so context doesn't degrade with position in the same way memory does for humans. However, models empirically perform worse on information buried in the middle of very long contexts (the "lost in the middle" phenomenon).
Training & Parameters

Training, Parameters & Backpropagation — Teaching a Model to Predict

Before a model can predict the next token, it has to learn from an almost incomprehensible volume of text. Pre-training a frontier LLM involves processing trillions of tokens scraped from the internet, books, code repositories, scientific papers, and countless other sources. But what exactly is "learning" in this context? At its core, it's adjusting numbers. An LLM is, beneath all its sophistication, a massive nested collection of numbers called parameters (also called weights). GPT-3 has 175 billion of them. GPT-4 likely has over a trillion. These numbers determine how the model transforms an input sequence of tokens into output probability distributions.

Training works like this: take a sequence of text from the training corpus. Feed the first N tokens into the model as input. Have the model predict the (N+1)th token. Compare its prediction to the actual (N+1)th token. Compute the error — a measure called the loss (typically cross-entropy loss). Then use an algorithm called backpropagation to figure out exactly how each of those hundreds of billions of parameters contributed to that error. Then nudge each parameter slightly in the direction that would have reduced the error. Repeat this process for trillions of token predictions, across thousands of GPU hours, and gradually the parameters settle into values that make the model remarkably good at predicting human-written text.

The scale is staggering. Training GPT-4 cost an estimated $100 million in compute. Training runs occupy thousands of NVIDIA A100 or H100 GPUs for months. The datasets contain more text than any human could read in thousands of lifetimes. And through this process of minimizing prediction error across that ocean of text, something extraordinary happens: the model doesn't just learn to predict tokens — it learns about the world, about logic, about language structure, about human values and conflicts and humor. All of it falls out of the single objective of predicting the next token accurately, at scale.

Real-World Analogy

Imagine teaching someone to be a great chef purely by having them taste millions of dishes and predict the next ingredient in the recipe. They'd taste, guess wrong, get corrected, taste again. Billions of guesses later, they wouldn't just know recipes — they'd understand flavor principles, cultural cuisine traditions, nutritional science, and the psychology of why certain combinations satisfy. Backpropagation is the "you got that wrong, here's the correction" signal. Parameters are the chef's accumulated intuition. The training corpus is the lifetime of meals.

LLM backpropagation training diagram showing forward pass prediction, loss calculation, and gradient flow back through neural network layers
# Simplified training loop — the essence of LLM pre-training
import torch
from torch.optim import AdamW

model = MyLLM(vocab_size=50257, d_model=768, n_heads=12, n_layers=12)
optimizer = AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)

for batch in dataloader:
 input_ids = batch["input_ids"] # [batch, seq_len]
 labels = batch["labels"] # [batch, seq_len] — shifted right

 # Forward pass: model predicts next token logits
 logits = model(input_ids) # [batch, seq_len, vocab_size]

 # Compute prediction error (cross-entropy loss)
 loss = F.cross_entropy(
 logits.view(-1, vocab_size),
 labels.view(-1)
 )

 # Backpropagation: compute gradients for all parameters
 loss.backward()

 # Gradient clipping — prevents explosions during training
 torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)

 # Update parameters in direction that reduces loss
 optimizer.step()
 optimizer.zero_grad()
⚡ Pro Tips — Training & Parameters
  • Parameter count ≠ capability. A well-trained smaller model (like Mistral 7B) often outperforms a poorly-trained larger one on specific tasks. Quality and curation of training data matters as much as model size — often more. This is why fine-tuning on high-quality domain-specific data produces such disproportionate results.
  • Perplexity is the training loss metric you'll see most often. Lower perplexity means the model assigns higher probability to the correct next token on average. It's a proxy for how well the model has "understood" its training corpus, but it doesn't directly measure capability on downstream tasks — that requires task-specific benchmarks.
  • You almost never need to pre-train from scratch. Use pre-trained models (GPT-4, Llama, Mistral) and either fine-tune them or use prompt engineering. Full pre-training is for companies with deep pockets and a specific need for custom training data distributions.
RLHF

Reinforcement Learning from Human Feedback — Making Models Actually Useful

A model trained purely on next token prediction produces something powerful but unreliable. It will complete sentences eloquently but will also hallucinate facts, produce harmful content, refuse no requests it shouldn't, and follow no requests it should. Pre-training teaches the model language and knowledge — but it doesn't teach it to be a helpful, honest, and harmless assistant. That alignment between model behavior and human values requires a separate training phase called Reinforcement Learning from Human Feedback (RLHF).

RLHF works in stages. First, human raters compare pairs of model outputs for the same prompt and indicate which response is better — more helpful, more accurate, more appropriate. These comparisons train a separate neural network called a reward model, which learns to predict human preference scores for any given response. Then, the main LLM is fine-tuned using reinforcement learning (specifically an algorithm called PPO — Proximal Policy Optimization) to produce responses that maximize the reward model's score. The model learns, through iterative training, to generate outputs that humans consistently rate as better.

RLHF is why ChatGPT feels so different from GPT-3 base model outputs. The same underlying architecture and similar parameters, but RLHF-tuned models are dramatically more helpful, more willing to admit uncertainty, better at following multi-step instructions, and more aligned with what users actually want. It's also why these models sometimes over-refuse — being too cautious can score well with raters who value safety, creating models that decline benign requests. The balance between helpfulness and safety in RLHF is one of the most active and difficult areas of AI alignment research.

Real-World Analogy

Pre-training is like producing a wildly talented but completely unrefined musician through sheer practice — technically extraordinary, no social skills, no sense of what an audience wants. RLHF is the music school that teaches them performance: how to read the room, when to play softly, how to engage an audience, what to skip. The talent was always there; RLHF shapes it into something an actual audience can appreciate and trust.

RLHF reinforcement learning from human feedback diagram showing human raters comparing outputs, reward model training, and PPO policy optimization loop
# Conceptual RLHF pipeline — simplified for clarity

# Step 1: Collect human preference data
preference_data = [
 {"prompt": "Explain recursion",
 "chosen": "Recursion is when a function calls itself...", # preferred
 "rejected": "Recursion = function calling function recursively"} # not preferred
]

# Step 2: Train a reward model on human preferences
reward_model = RewardModel.from_pretrained("gpt2")
# reward_model learns to score: chosen > rejected

# Step 3: Use PPO to maximize reward while staying close to base policy
from trl import PPOTrainer, PPOConfig

ppo_config = PPOConfig(
 learning_rate=1.4e-5,
 batch_size=64,
 kl_penalty="kl", # prevents model drifting too far from base
)
trainer = PPOTrainer(config=ppo_config, model=policy_model,
 ref_model=base_model, reward_model=reward_model)

# Each PPO step: generate response → score with reward model → update policy
trainer.step(queries, responses)
⚡ Pro Tips — RLHF and Alignment
  • DPO (Direct Preference Optimization) is increasingly replacing PPO-based RLHF. It directly optimizes the LLM on preference pairs without a separate reward model, making it simpler, more stable, and less compute-intensive. If you're fine-tuning for alignment, DPO is now the practical default.
  • RLHF can introduce reward hacking: the model learns to maximize reward model scores in ways that don't actually correspond to better human satisfaction. Long, verbose responses often score higher even when concise answers are better. Watch for verbose drift in RLHF-trained models.
  • Constitutional AI (Anthropic's approach) and RLAIF (RL from AI Feedback) are alternatives/complements to human-labeled RLHF — using an AI model to generate and evaluate preference data at a fraction of the cost. These are becoming increasingly important as the bottleneck shifts from compute to human annotation capacity.
Transformer Architecture

The Transformer Architecture — The Revolution That Changed Everything

Before 2017, language models processed text sequentially — word by word, hidden state to hidden state. Recurrent Neural Networks (RNNs) and Long Short-Term Memory networks (LSTMs) were state of the art. They were slow (each token depended on the previous computation), struggled with long-range dependencies (information from 500 tokens ago was diluted by the time it reached the current position), and were nearly impossible to parallelize efficiently. Then a team of Google researchers published "Attention Is All You Need" and introduced the Transformer. Everything changed.

The Transformer's revolutionary insight: process the entire input sequence in parallel, not sequentially. Instead of passing hidden states from token to token, the Transformer lets every token directly attend to every other token in the sequence simultaneously. This makes training massively parallelizable (perfect for GPU/TPU clusters), eliminates the vanishing gradient problem that plagued RNNs, and allows the model to capture relationships between tokens regardless of how far apart they appear in the sequence. A word at position 1 can directly influence the processing of a word at position 10,000 without any intermediate steps.

The Transformer consists of stacked layers (GPT-3 has 96 layers), each containing two main components: a multi-head attention mechanism and a feed-forward neural network. The attention mechanism contextualizes each token's representation based on all other tokens (more on this in the next section). The feed-forward network then applies the same learned transformation independently to each position — this is where much of the model's "knowledge" is stored as learned patterns in the weights. Both components use residual connections and layer normalization, which make very deep networks stable to train.

Transformer architecture diagram showing stacked encoder-decoder layers with multi-head attention, feed-forward networks and residual connections
import torch
import torch.nn as nn

class TransformerBlock(nn.Module):
 """One Transformer layer: attention + feed-forward + residual connections"""
 def __init__(self, d_model: int, n_heads: int, d_ff: int, dropout: float = 0.1):
 super().__init__()
 # Multi-head self-attention
 self.attn = nn.MultiheadAttention(d_model, n_heads, dropout=dropout, batch_first=True)
 self.norm1 = nn.LayerNorm(d_model)
 # Feed-forward network (where knowledge is stored)
 self.ff = nn.Sequential(
 nn.Linear(d_model, d_ff),
 nn.GELU(), # modern activation function
 nn.Dropout(dropout),
 nn.Linear(d_ff, d_model),
 )
 self.norm2 = nn.LayerNorm(d_model)
 self.drop = nn.Dropout(dropout)

 def forward(self, x, mask=None):
 # Attention with residual connection (x + attention output)
 attn_out, _ = self.attn(x, x, x, attn_mask=mask)
 x = self.norm1(x + self.drop(attn_out)) # residual + norm
 # Feed-forward with residual connection
 x = self.norm2(x + self.drop(self.ff(x))) # residual + norm
 return x
⚡ Pro Tips — Transformer Architecture
  • The quadratic scaling problem: attention between N tokens requires O(N²) computation and memory. A 1,000-token sequence requires 1M attention computations; 100,000 tokens requires 10 billion. This is why context window expansion (to 128k, 1M tokens) is a significant engineering challenge, requiring techniques like Flash Attention, sparse attention, and ring attention.
  • Positional encoding is how the Transformer knows token order — since attention is permutation-invariant (it doesn't inherently know which token comes first). Modern models use Rotary Position Embeddings (RoPE) which encode relative positions and generalize better to sequence lengths beyond what was seen in training.
  • The feed-forward layer is typically 4× the model dimension. So a 768-dimensional model has a 3072-dimensional feed-forward layer. This is where most parameters live and where most model "knowledge" is believed to be stored (based on interpretability research). When you're choosing model size, the feed-forward expansion ratio matters as much as the number of layers.
Attention Mechanism

Attention — How Every Token Listens to Every Other Token

The attention mechanism is the single most important conceptual breakthrough in modern AI. Understanding it at a technical level transforms your intuition for why LLMs behave the way they do. Here's the core idea: for each token in the sequence, attention produces a new, context-aware representation by computing a weighted sum of all other tokens' representations — with weights determined by how "relevant" each other token is to the current one. The word "bank" in "I walked to the river bank" gets a very different internal representation than in "I deposited money at the bank" — because the other words in the sequence provide different context weights.

Mechanically, attention computes three matrices from each token's representation: a Query (Q, "what am I looking for?"), a Key (K, "what do I offer?"), and a Value (V, "what information do I carry?"). The attention score between two tokens is the dot product of one's Query and the other's Key — scaled by the square root of the dimension for numerical stability — passed through softmax to produce probabilities. These probabilities are then used to weight-average the Value vectors, producing an output that blends information from all relevant tokens. The whole thing is differentiable and trains end-to-end with backpropagation.

Multi-head attention runs this process multiple times in parallel (typically 8, 12, 32, or 96 heads depending on model size), with different Q/K/V projection matrices for each head. Each head can specialize in attending to different types of relationships — one head might track syntactic subject-verb agreement, another might track co-reference (which pronoun refers to which noun), another might capture long-range semantic dependencies. The heads' outputs are concatenated and projected back to the model dimension. This parallel specialization is what gives attention its extraordinary expressiveness.

Multi-head attention mechanism diagram showing Query Key Value matrices, scaled dot-product attention scores and multiple parallel attention heads
# Scaled dot-product attention — the mathematical heart of the Transformer
import torch
import torch.nn.functional as F
import math

def scaled_dot_product_attention(Q, K, V, mask=None):
 """
 Q, K, V: [batch, heads, seq_len, d_head]
 Returns: weighted sum of values, attention weights
 """
 d_k = Q.size(-1) # dimension of each head

 # Step 1: Dot product of Query and Key (how similar?)
 scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(d_k)
 # scores shape: [batch, heads, seq_len, seq_len]

 # Step 2: Apply causal mask (can't attend to future tokens)
 if mask is not None:
 scores = scores.masked_fill(mask == 0, float("-inf"))

 # Step 3: Softmax — convert scores to attention probabilities
 attn_weights = F.softmax(scores, dim=-1)

 # Step 4: Weighted sum of Values
 output = torch.matmul(attn_weights, V)
 return output, attn_weights
⚡ Pro Tips — Attention
  • The causal mask in decoder-only models (like GPT) is critical: during training, the model must not attend to future tokens, otherwise it would just copy the answer rather than predict it. This lower-triangular mask ensures each position can only attend to previous positions and itself.
  • Flash Attention (Dao et al., 2022) is a hardware-aware attention implementation that reduces memory usage from O(N²) to O(N) by computing attention in tiles that fit in GPU SRAM, avoiding slow HBM (high-bandwidth memory) reads. It's 2–4× faster than standard attention and is now standard in all production frameworks. Always use it.
  • Grouped Query Attention (GQA), used in Llama 2/3 and other models, reduces the number of K/V heads relative to Q heads — dramatically cutting memory during inference (KV cache) without significant capability loss. This is why smaller models can now serve longer contexts efficiently.
Emergent Behavior

Emergent Behavior — The Mystery at the Center of Modern AI

Here we reach the most honest and humbling part of this entire story: neither the researchers who build these systems nor the engineers who deploy them fully understand why they work as well as they do. The architecture is carefully designed. The training objectives are clearly specified. The mathematical operations are precisely defined. And yet the resulting capabilities — multi-step reasoning, code generation, creative writing, language translation between languages the model never explicitly paired — were not directly designed. They emerged from the scale and structure of the system in ways that surprised their creators.

Emergent behaviors in LLMs are capabilities that appear abruptly at certain scales rather than improving smoothly with model size. Researchers documented this phenomenon extensively: a model at 7B parameters might be barely better than random at 5-shot arithmetic. The same architecture at 70B parameters might solve the same problems reliably. At 100B+, it might solve them and explain its reasoning. These phase transitions — where capability appears to jump qualitatively rather than improve gradually — suggest that something complex is happening in the learned representations that isn't captured by simple metrics.

Chain-of-thought reasoning is the most striking example. Prompting a model to "think step by step" before answering dramatically improves accuracy on multi-step reasoning tasks. Nobody explicitly trained for this behavior by writing "reason step by step" examples in the pre-training data (or at least not primarily). It emerged from the model's learned statistical patterns about how humans structure their reasoning in text. The model has learned, from reading millions of solved problems, that showing work correlates with correct answers — and can apply this pattern to new problems.

Real-World Analogy

Picture a city that was planned with roads and zoning but no design for culture. Nobody specified "there will be jazz clubs here" or "this neighborhood will become an art district." But given enough people, infrastructure, and economic activity, cultural scenes emerge. They're not random — they follow from the underlying conditions — but they weren't directly planned. LLM emergent behaviors are like these cultural scenes: real, useful, structurally motivated, and not explicitly designed by anyone.

LLM emergent behavior chart showing sudden capability jumps at scale with phase transitions in reasoning arithmetic and language tasks ⚡ Pro Tips — Working with Emergent Behavior
  • Emergent capabilities are scale-dependent. Don't assume a capability you observed in GPT-4 exists in a smaller, cheaper model you're deploying. Always benchmark the specific model on your specific task. Model capability rankings change across task types — a model excellent at code may underperform on structured data extraction.
  • Chain-of-thought prompting is your most reliable lever for improving complex reasoning. "Let's think step by step" before the answer, or asking the model to show its work, consistently improves accuracy on multi-step problems. This works even with API calls — include it in your system prompt for reasoning-intensive tasks.
  • The fact that we can't fully explain emergent behavior is a feature of honest science, not a reason for mistrust. We can't fully explain why certain drug combinations work, but we use them based on empirical evidence. Evaluate LLMs empirically on your use case. Don't assume capabilities; measure them.
Synthesis

How It All Connects — The Complete LLM Mental Model

Let's trace the complete journey, from a pre-training dataset to your ChatGPT conversation. Trillions of tokens of internet text, books, and code are tokenized and fed through the training pipeline. A Transformer architecture — stacked blocks of multi-head attention and feed-forward networks — processes these tokens in parallel, with each layer enriching each token's representation with context from every other token via attention. Backpropagation adjusts hundreds of billions of parameters to minimize prediction error on the training data. After months of training on thousands of GPUs, you have a base model that can generate fluent text but isn't necessarily helpful or safe.

RLHF then fine-tunes this base model against human preference data — human raters comparing pairs of outputs, a reward model learning to predict those preferences, and PPO or DPO optimization pushing the model toward higher-reward outputs. The result is the chat model you interact with: architecturally identical to the base model, behaviorally dramatically different. When you send a message, it's tokenized, embedded as numerical vectors, and passed through the full stack of Transformer layers. Each layer refines the representations. The final layer produces logits over the vocabulary. Temperature and sampling parameters determine which token is selected. That token is appended and the process repeats. Token by token, your response assembles from a system that learned everything it knows by predicting the next word, at staggering scale, across all of human writing.

Getting Started

Getting Started — Run Your First LLM in 5 Minutes

# Option A: Use the Anthropic API (Claude)
pip install anthropic

# Option B: Run a local LLM with Ollama (no API key needed)
# Install Ollama: https://ollama.ai
ollama pull llama3.2 # ~2GB download
ollama run llama3.2 # interactive chat in your terminal

# Option C: HuggingFace Transformers (full control)
pip install transformers torch
# Complete working example — token-by-token generation with HuggingFace
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

# Load a small but capable model
model_name = "microsoft/phi-2" # 2.7B params, runs on consumer hardware
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
 model_name, torch_dtype=torch.float16, device_map="auto"
)

# Tokenize input
prompt = "Explain why the Transformer architecture was a breakthrough:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

# Generate — token by token under the hood
with torch.no_grad():
 output_ids = model.generate(
 **inputs,
 max_new_tokens=200,
 temperature=0.7,
 do_sample=True,
 top_p=0.9,
 pad_token_id=tokenizer.eos_token_id
 )

# Decode back to text
response = tokenizer.decode(output_ids[0], skip_special_tokens=True)
print(response)
FAQ

Frequently Asked Questions

How do large language models actually work? LLMs are neural networks trained to predict the next token in a sequence. The Transformer architecture processes all tokens in parallel using attention mechanisms, which let each token build a context-aware representation by attending to all other tokens. With hundreds of billions of parameters tuned via backpropagation on trillions of tokens, the model learns to predict human text so accurately that coherent language and reasoning emerge as byproducts. What are parameters in an LLM? Parameters (also called weights) are the adjustable numerical values in the neural network. An LLM like GPT-4 has hundreds of billions to over a trillion parameters. During training, backpropagation adjusts these values to minimize prediction error. After training, the parameters are frozen — they encode everything the model "knows" about language, facts, reasoning patterns, and world knowledge learned from the training corpus. What is RLHF and why does it matter? Reinforcement Learning from Human Feedback (RLHF) is the process that turns a raw pre-trained language model into a helpful, safe assistant. Human raters compare model outputs and indicate preferences. A reward model learns to predict these preferences. The main LLM is then fine-tuned (using PPO or DPO) to produce outputs that maximize reward scores. RLHF is why ChatGPT-style models feel so different from raw language model outputs — it aligns the model's behavior with human values and expectations. What is the Transformer architecture? The Transformer is a neural network architecture introduced in 2017 that processes entire sequences in parallel rather than word-by-word. It uses attention mechanisms to let every token directly attend to every other token, capturing long-range dependencies without information loss. Stacked Transformer blocks, each containing multi-head attention and feed-forward networks with residual connections, build increasingly rich representations. All modern LLMs are variants of this architecture. What is attention in an LLM? Attention is the mechanism that gives each token a context-aware representation by computing a weighted average of all other tokens' information. It works via Query, Key, and Value matrices: the similarity between a token's Query and other tokens' Keys determines attention weights, and those weights determine how much of each token's Value information flows into the current token's representation. Multi-head attention runs this process multiple times in parallel, with each head potentially specializing in different linguistic relationships. What is emergent behavior in LLMs? Emergent behaviors are capabilities that appear abruptly at certain model scales rather than improving smoothly with size. Examples include chain-of-thought reasoning, arithmetic, code generation, and in-context learning. These weren't directly programmed — they emerged from training large models on vast data to predict text. Even the researchers who built these systems can't fully explain all emergent capabilities, making it one of the most active areas of AI interpretability research. How is a language model different from a chatbot? A language model is the underlying neural network that predicts next tokens. A chatbot is the product built on top of it — typically a base language model that has been fine-tuned with RLHF, given a system prompt, and wrapped in a conversation interface. ChatGPT, Claude, and Gemini are chatbots; GPT-4, Claude 3.5, and Gemini 1.5 Pro are the underlying language models. Base models can do many things chatbots can't (like continuing a story mid-sentence), and chatbots add safety alignment and instruction-following that base models lack. Can I run an LLM locally on my own hardware? Yes — tools like Ollama, LM Studio, and GPT4All make this accessible. Models like Llama 3.2 (3B), Phi-3 Mini (3.8B), and Mistral 7B run on modern laptops with 16GB RAM using 4-bit quantization. For production-quality models comparable to GPT-4, you'll need significant GPU resources (A100 80GB or multiple GPUs) — but for experimentation and many production use cases, 7B–13B quantized models are surprisingly capable.

🔬 LLM Mechanics Interactive Lab

Watch next-token prediction in real time, visualize attention weights across sentences, simulate backpropagation training, rate RLHF responses, and explore emergent behaviors.

Next Token Prediction — Live

Watch the model predict one token at a time. See the probability distribution over candidate tokens before each selection.

Token Stream Step 0 Top candidate tokens — probability distribution

Attention Weights — Interactive Visualization

Click any word in the table to highlight which other words it pays the most attention to. Darker cells = stronger attention.

Attention Heatmap Click a word to see its attention pattern Click a word to see which tokens it attends to most strongly.

Backpropagation — Watch Parameters Update

Simulate a training step: forward pass predicts a token, loss is computed, gradients flow backward, and parameters update.

Neural Network — Training Step Idle Training Loss (lower = better) LowLoss: 3.2High Training Log

RLHF — Be a Human Rater

Compare two model responses to the same prompt. Pick the better one. This is exactly what human raters do to train the reward model in RLHF.

Human Preference Collection Score: 0/0 Prompt

Emergent Behaviors — Not Explicitly Trained

These capabilities weren't directly designed — they emerged from training at scale. Click each to understand what it is and why it's remarkable.

Emergent Capability Explorer Click any emergent behavior to learn what it is and why it matters.

Knowledge Check — Test Your LLM Understanding

8 questions covering all core LLM concepts.

Tags
LLMlarge language modelstransformer architectureattention mechanismRLHFbackpropagationGPTemergent behaviorAI engineeringneural networks
Share this article