LLMAI parameterstemperatureTop-PTop-Kstop sequencesfrequency penaltypresence penaltyAI engineeringlanguage models
TL;DR Temperature, Top-P, Top-K, Stop Sequences, Frequency Penalty, Presence Penalty — these aren't optional knobs. They're the difference between an AI that rambles and one that performs like a precision instrument. Master them here with real code, real analogies, and an interactive playground.
You've called the API. You got an output. But it's either too robotic, too random, too repetitive, or it won't stop talking. The fix isn't a better prompt — it's understanding the six parameters that sit between your prompt and the model's output. This guide covers all of them, deeply.
Temperature Top-P Top-K Stop Sequences Frequency Penalty Presence Penalty Post ExcerptTemperature, Top-P, Top-K, Stop Sequences, Frequency Penalty, Presence Penalty — these aren't optional knobs. They're the difference between an AI that rambles and one that performs like a precision instrument. Master them here with real code, real analogies, and an interactive playground.
IntroductionPicture this: you've spent two hours crafting the perfect system prompt. It's detailed, structured, and full of examples. You hit send, and the model spits back something technically correct but completely off — either a one-word answer when you needed three paragraphs, or a creative novella when you asked for a customer service reply. You tweak the prompt again. And again. Sound familiar?
Here's the thing most tutorials miss: prompts tell the model what to do. Parameters control how it thinks. The probability distribution over the model's entire vocabulary gets filtered, weighted, and trimmed by six specific numbers before a single token is chosen. Miss those numbers, and no prompt will save you.
LLMs are, at their core, probability machines. For every position in the output, the model produces a giant vector of probabilities — one per token in its vocabulary (often 50,000+ tokens). Before any token gets selected, these six parameters reshape that probability distribution. Temperature flattens or sharpens it. Top-P and Top-K trim it. Frequency and Presence penalties dent it. Stop sequences abort it entirely. That's the whole game.
Counterintuitive InsightMore temperature ≠ more intelligent. A common mistake: developers crank up temperature thinking it makes the model "think harder." It doesn't. Higher temperature just means the model is more willing to pick lower-probability tokens — which can mean creativity, but also hallucination, incoherence, and factual errors. Intelligence is baked into the weights at training time. Temperature only affects sampling randomness.
TemperatureTemperature is the most-discussed LLM parameter for good reason: it's intuitive, powerful, and wildly misunderstood. The parameter typically ranges from 0 to 2, though most production use-cases stay between 0 and 1. At its technical core, temperature is a divisor applied to the model's raw logits (the pre-softmax scores) before the probability distribution is computed. A lower temperature amplifies the differences between scores, making high-probability tokens even more dominant. A higher temperature compresses the differences, giving lower-probability tokens a fighting chance.
At temperature = 0, you're asking the model to always pick the single most probable token — deterministic, predictable, zero creative wiggle room. At temperature = 1, you're sampling directly from the model's raw distribution. Above 1.0, you're actively flattening the distribution to inject randomness beyond what the training data would suggest. For most production agents, the sweet spot is 0.1–0.3 for reliability, 0.7–1.0 for creativity.
Real-World AnalogyImagine a drive-thru ordering agent. "The customer said large fries" should always produce "Added large fries to your order" — not "Certainly! I've joyfully registered your magnificent fry selection." That's temperature near zero. Now imagine a lyric-writing assistant — you want it to surprise you, pick unexpected metaphors, break from cliché. That's temperature near 1. Same model, same weights, completely different behavior.
import anthropic
client = anthropic.Anthropic()
# Low temperature — consistent, predictable (customer service bot)
response_precise = client.messages.create(
model="claude-opus-4-6",
max_tokens=256,
temperature=0.1, # near-deterministic
messages=[{"role": "user", "content": "What are your store hours?"}]
)
# High temperature — varied, creative (poem generator)
response_creative = client.messages.create(
model="claude-opus-4-6",
max_tokens=512,
temperature=0.9, # embrace variation
messages=[{"role": "user", "content": "Write a poem about autumn rain"}]
)
⚡ Pro Tips / Common Mistakes — Temperature
Top-P, also called Nucleus Sampling, takes a more elegant approach than its cousin Top-K. Instead of asking "how many tokens should I consider?", it asks "what's the smallest set of tokens whose combined probability exceeds P?" If P = 0.9, the model builds a shortlist by walking down the ranked token list — most probable first — until the running total hits 90%. It then samples exclusively from that shortlist.
Why is this smarter than Top-K? Because probability distributions are context-dependent. When continuing "The dog sat on the…", maybe only 5 tokens (mat, floor, couch, grass, chair) account for 90% of probability — Top-P naturally constrains to those. But when continuing "She felt a strange sense of…", hundreds of semantically valid continuations exist, and Top-P expands to include them. Top-K would apply the same hard number in both cases, either over-constraining the first scenario or under-constraining the second.
Real-World AnalogyThink of Top-P as a restaurant menu that adjusts based on how many good options exist. If you ask for "a great Italian dish" at a specialized trattoria, the nucleus is small — three or four standout options. If you ask at a global fusion restaurant, the nucleus expands. Same threshold (90% of what's good), different pool size. Top-K would just say "give me 10 options" regardless of whether the 10th option is excellent or garbage.
# Top-P = 0.9 — safe default, good for most tasks
response = client.messages.create(
model="claude-opus-4-6",
max_tokens=400,
temperature=1.0,
top_p=0.9, # only sample from top 90% probability mass
messages=[{"role": "user", "content": "Continue this story: The lighthouse keeper"}]
)
# Top-P = 0.5 — conservative, higher quality tokens only
response_tight = client.messages.create(
model="claude-opus-4-6",
max_tokens=400,
top_p=0.5, # very focused nucleus — only highest-confidence tokens
messages=[{"role": "user", "content": "Summarize this legal document"}]
)
⚡ Pro Tips / Common Mistakes — Top-P
Top-K is the blunter tool of the sampling family. Set Top-K = 40, and the model considers only the 40 most probable tokens at each step — regardless of whether number 40 has a 15% probability or a 0.001% probability. The rest of the vocabulary is zeroed out completely. It's a hard wall, not a soft threshold. This makes it faster to compute and easier to reason about, which is why it remains popular in production environments even as Top-P has largely superseded it for quality.
The counterintuitive failure mode of Top-K: at very low values (K = 5 or 10), you get repetitive and predictable text — not because the model is being careful, but because you've removed the diversity that makes language feel natural. At very high values (K = 500+), Top-K becomes essentially ineffective since you're including tokens with negligible probability. The sweet spot is typically K = 40–80 for general use, though many practitioners find Top-P alone performs better across diverse tasks.
# Top-K with OpenAI-compatible API
import openai
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Write a product description"}],
# Note: OpenAI doesn't expose Top-K directly
# Use Hugging Face / local models for Top-K control
)
# Hugging Face Transformers — full control
from transformers import pipeline
gen = pipeline("text-generation", model="mistralai/Mistral-7B-Instruct-v0.2")
output = gen(
"The future of AI is",
max_new_tokens=100,
do_sample=True,
top_k=50, # consider only top 50 tokens
top_p=0.92, # can stack with top_p
temperature=0.8
)
⚡ Pro Tips / Common Mistakes — Top-K
LLMs don't have an inherent sense of "done." Left to their own devices, they'll continue generating tokens until they hit max_tokens, sometimes producing beautifully structured output, sometimes drifting into nonsense, and sometimes — critically for agents — inventing the other side of a conversation. Stop sequences are the circuit breaker. You define one or more string patterns, and generation halts the moment any of them appear.
This is essential for structured output and conversation control. Building a chat agent where the model plays the role of a support rep? You'd set "\nCustomer:" as a stop sequence. The model generates the agent's response and the instant it starts to write what the customer would say next (which it will, given enough rope), generation cuts off cleanly. You're not hoping the model knows to stop — you're enforcing it programmatically.
Think of stop sequences like the director yelling "cut!" on a film set. The actor (the model) knows their lines and could improvise indefinitely. But the director defines exactly where the scene ends. Without a cut, you'd get beautiful improv that runs three hours over schedule and blows your budget. Stop sequences are how you get the exact scene length you need, consistently, every take.
# Stop sequence for a customer service agent
response = client.messages.create(
model="claude-opus-4-6",
max_tokens=512,
stop_sequences=["\nCustomer:", "\nUser:", "###"],
system="You are a support agent. Respond only as Agent:",
messages=[{
"role": "user",
"content": "Agent: Hello! How can I help you today?\nCustomer: My order is late."
}]
)
# Stop on JSON completion
response_json = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024,
stop_sequences=["}"], # stop after closing brace (careful — use wisely)
messages=[{"role": "user", "content": 'Output a JSON object: {"name": ...'}]
)
⚡ Pro Tips / Common Mistakes — Stop Sequences
"\n\n###", role labels, or XML-style tags.If you've ever asked an LLM to write a long-form article and noticed it using the phrase "it's worth noting" seventeen times, you've encountered the repetition problem. Models trained on human text learn that certain phrases are common — and they overuse them proportionally. The Frequency Penalty is a taxation system: every time a token appears in the output so far, its probability gets reduced by a fixed amount (typically the penalty value × the token's count). The more a token repeats, the more heavily it's penalized on subsequent selections.
This is proportional penalization, which means the first occurrence is free, the second occurrence pays a small tax, the third occurrence pays more. Values typically range from 0 to 2. A value of 0.3–0.6 noticeably reduces repetition without making the output feel artificially varied. Values above 1.0 can cause the model to avoid important domain-specific terms that genuinely need repeating (like a product name in a marketing brief).
# OpenAI — frequency_penalty for long-form content
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Write a 500-word blog intro about machine learning"}],
max_tokens=600,
frequency_penalty=0.4, # reduce phrase repetition
presence_penalty=0.0 # no presence penalty here
)
# High frequency_penalty use case: diverse vocabulary for marketing copy
response_rich = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Write 5 unique taglines for a coffee brand"}],
max_tokens=200,
frequency_penalty=0.8 # strongly discourage word reuse
)
⚡ Pro Tips / Common Mistakes — Frequency Penalty
Presence Penalty is the binary twin of Frequency Penalty. While Frequency Penalty scales its punishment based on how many times a token has appeared, Presence Penalty applies a flat tax to any token that has appeared at least once — regardless of whether it appeared one time or twenty. The result is a model that actively explores new vocabulary rather than doubling down on words it's already used.
Think of it as the difference between a tax system and a ban. Frequency Penalty taxes each additional use at an increasing rate. Presence Penalty just says "you used that word — it's now slightly less attractive for the rest of this generation." In practice, Presence Penalty tends to produce outputs with broader topic coverage and more diverse word choices — useful for brainstorming, ideation, or generating content that should feel fresh across multiple sections. Combined with a moderate Frequency Penalty, you get both reduced repetition and increased novelty.
# Presence penalty for brainstorming — maximum topic diversity
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "user",
"content": "List 15 unique business ideas in the health tech space"
}],
max_tokens=800,
frequency_penalty=0.3, # reduce phrase repetition
presence_penalty=0.6 # push toward new concepts each sentence
)
# Zero both for highly constrained/legal text
response_legal = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Draft a Terms of Service clause"}],
max_tokens=400,
frequency_penalty=0.0, # legal text needs precise repetition
presence_penalty=0.0
)
⚡ Pro Tips / Common Mistakes — Presence Penalty
These six parameters don't operate in isolation. Every time the model selects a token, they all fire in sequence: the raw logits come out of the neural network → Frequency and Presence Penalties adjust the logit scores based on what's already been generated → Temperature scales the adjusted logits → Top-K trims the vocabulary to the top K candidates → Top-P further trims to the nucleus → one token is sampled from the remaining distribution → Stop Sequences check if the output should halt.
This sequential application means parameters can amplify or cancel each other. A high Frequency Penalty with a low Temperature = diverse word choices but very predictable sentence structure. A high Top-P with high Temperature = genuinely random-feeling text with a broad vocabulary. Zero everything except a tight Stop Sequence = the model generates its best guess and you control exactly when it stops. There's no universal "best config" — there's only the right config for your task.
The Config Matrix — What to Use WhenFor a customer service chatbot: Temperature 0.2, Top-P 0.9, Frequency Penalty 0.2, Stop Sequences set to role boundaries. For a code assistant: Temperature 0.1, Top-P 0.95, both penalties at 0. For a creative writing tool: Temperature 0.85, Top-P 0.92, Frequency Penalty 0.4, Presence Penalty 0.3. For a data extraction agent: Temperature 0, Stop Sequences on your schema delimiter, penalties at 0. Context is everything.
Getting StartedRather than theory, here are copy-paste starting configurations you can drop into your projects. Each is annotated with the reasoning behind each choice. Treat them as baselines — measure your output quality, then tune from there.
# ── CONFIG TEMPLATES ──────────────────────────────────────────
CONFIGS = {
# 1. Customer Support / FAQ Bot
"support_bot": {
"temperature": 0.2, # consistent, factual responses
"top_p": 0.9, # some nucleus flexibility
"frequency_penalty": 0.3, # avoid filler repetition
"presence_penalty": 0.0,
"stop_sequences": ["\nCustomer:", "\nUser:"],
"max_tokens": 300
},
# 2. Creative Writing Assistant
"creative_writer": {
"temperature": 0.85,
"top_p": 0.92,
"frequency_penalty": 0.45, # varied vocabulary
"presence_penalty": 0.3, # explore new concepts
"stop_sequences": [],
"max_tokens": 1024
},
# 3. Code Generation
"code_gen": {
"temperature": 0.1, # near-deterministic
"top_p": 0.95,
"frequency_penalty": 0.0, # code repeats identifiers — that's correct
"presence_penalty": 0.0,
"stop_sequences": ["```"], # stop at code block end
"max_tokens": 2048
},
# 4. Data Extraction Agent
"data_extractor": {
"temperature": 0.0, # fully deterministic
"top_p": 1.0, # irrelevant at temp=0
"frequency_penalty": 0.0,
"presence_penalty": 0.0,
"stop_sequences": ["}", ""],
"max_tokens": 512
},
# 5. Brainstorming / Ideation
"brainstorm": {
"temperature": 1.0,
"top_p": 0.95,
"frequency_penalty": 0.6,
"presence_penalty": 0.7, # maximum novelty
"stop_sequences": [],
"max_tokens": 600
}
}
# Usage
def call_model(prompt, config_name="support_bot"):
cfg = CONFIGS[config_name]
return openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
**cfg
)
FAQ
Adjust all 6 parameters live. See exactly how each one shapes the AI's output — with real token probability visualization and side-by-side comparisons.
Scenario Preset temp: 0.85 top_p: 0.92 top_k: 50 freq_penalty: 0.45 pres_penalty: 0.30 stop: none 🌡 Temperature Controls randomness. Low = consistent, High = creative. 0.85 0 · Deterministic1 · Balanced2 · Wild ◎ Top-P (Nucleus) Probability mass threshold. Smaller = more focused. 0.92 0.01 · Tight0.5 · Moderate1.0 · All tokens NUCLEUS SIZE ▦ Top-K Hard limit on candidate tokens. 0 = disabled. 50 0 · Off50 · Balanced200 · Open ⊠ Stop Sequences Halt generation on these tokens/patterns. ♻ Frequency Penalty Tax on repeated tokens. Scales with frequency. 0.45 0 · None1 · Strong2 · Max ◈ Presence Penalty Binary tax on any token already used once. 0.30 0 · Off1 · Active2 · Aggressive Simulated Output ⚗️Configure parameters and click Generate
Token Probability Visualization Randomness Vocab Diversity Coherence Creativity Side-by-Side: Low vs High Temperature 🌡 Temperature: 0.1 (Deterministic) Click Generate to populate 🔥 Temperature: 1.0 (Creative) Click Generate to populate