Beginner Fundamentals
What is a large language model (LLM)?
An LLM is a model trained on huge amounts of text to predict the most likely next token given what came before. That training gives it broad, flexible language ability, but it's pattern-matching, not human-style reasoning — it can produce fluent, confident text that is nonetheless factually wrong.
What is prompt engineering?
Prompt engineering is the practice of designing and refining inputs to an AI model to reliably get the output you want — through specificity, examples, context, and structure — rather than hoping the model guesses your intent from a vague request.
What is a token, and why does it matter?
A token is a chunk of text, often smaller than a whole word, that a model processes as its basic unit. Token count directly determines how much fits in a model's context window and is commonly the basis for API pricing.
What is the difference between zero-shot and few-shot prompting?
Zero-shot prompting asks directly with no examples, relying on the model's general training. Few-shot prompting includes a small number of worked examples of the desired input/output pattern first, which improves consistency for tasks with a specific format or style.
What is an AI "hallucination"?
A hallucination is confidently stated information from a model that is actually incorrect or entirely fabricated — a citation that doesn't exist, a fact that was never true, a code API that isn't real. It happens because the model generates statistically plausible text with no built-in fact-checking mechanism.
What is a context window?
The context window is the maximum amount of text (measured in tokens) a model can consider at once — your prompt plus conversation history plus any pasted documents. Content beyond that limit has to be dropped or summarized.
Intermediate Techniques & Architecture
What is the difference between temperature and top-p sampling?
Temperature scales how much probability mass spreads across less-likely tokens — higher temperature means more randomness. Top-p (nucleus sampling) instead selects from the smallest set of tokens whose cumulative probability exceeds a threshold p, dynamically adjusting the candidate pool size rather than uniformly reshaping the whole distribution. The two are often used together in practice.
What is chain-of-thought prompting, and why does it improve accuracy?
Chain-of-thought prompting asks the model to reason through a problem step by step before giving a final answer, instead of jumping straight to a conclusion. Making the reasoning explicit gives the model a chance to catch its own errors along the way, which measurably improves accuracy on multi-step reasoning tasks.
What is RAG, and why is it used?
Retrieval-Augmented Generation (RAG) retrieves relevant chunks from an external knowledge source (like your own documents) and inserts them into the prompt as context before generation. This grounds answers in real, specific, or up-to-date source material instead of relying solely on the model's training data — commonly used to reduce hallucination on private or fast-changing information.
What is the difference between a system prompt and a user prompt?
A system prompt sets overall behavior and rules for the whole conversation, configured once and typically taking priority over conflicting instructions. User prompts are the ongoing back-and-forth messages in the actual conversation. Developers usually configure persistent behavior at the system level rather than repeating instructions in every user turn.
What is structured output / JSON mode, and why would you use it?
Structured output constrains a model's response to a strict, parseable format — commonly JSON matching a specified schema — so downstream code can reliably parse the result. Dedicated structured-output features enforce valid syntax more consistently than simply asking for JSON in plain text.
What is prompt injection?
Prompt injection is when untrusted text an AI system processes — embedded in a webpage, document, or email — contains hidden instructions designed to override the system's original task. It's a real security concern whenever an AI reads content it didn't originate, and a reason to keep human review in the loop for sensitive agent actions.
What is few-shot prompting most useful for, compared to zero-shot?
Few-shot prompting earns its extra length on tasks where output format or judgment consistency really matters — like classification with a specific label set, or matching a particular writing style across many generated items. For simple, common tasks, zero-shot is often sufficient.
Advanced Deep Dives
What is the difference between prompting, RAG, and fine-tuning, and how would you decide between them?
These are three ways to customize AI behavior, roughly in order of increasing cost and effort: prompting (better instructions/examples in the prompt itself — cheapest, fastest, no setup), RAG (grounding answers in your own documents via retrieval, without retraining), and fine-tuning (retraining the model on your own example data — most expensive and slowest). A reasonable decision process: start with prompting; reach for RAG when answers need grounding in specific/private/changing data prompting can't provide; reach for fine-tuning only when the other two genuinely can't achieve the needed behavior change.
Why can chain-of-thought prompting sometimes produce a confidently wrong answer despite showing "reasoning"?
The visible reasoning steps are still generated text, following learned patterns — they aren't a guarantee of genuine logical validity. A model can produce plausible-looking intermediate steps that don't actually follow from each other, or that rationalize a conclusion it would have reached anyway. Visible reasoning helps you audit the logic yourself, but it isn't proof of correctness on its own.
In a RAG system, what happens if the retrieval step returns irrelevant or low-quality context?
The generation step is only as good as what it's grounded in — irrelevant or low-quality retrieved context can lead the model to produce an answer that's confidently wrong or off-topic, sometimes worse than if it had used only its own training knowledge. This is why retrieval quality (chunking strategy, embedding model choice, similarity thresholds) is often the actual bottleneck in a RAG system's real-world accuracy, not the generation model itself.
What's a practical mitigation for prompt injection in an AI agent that can take real-world actions?
Common mitigations include: keeping a human in the loop for sensitive or irreversible actions rather than fully autonomous execution, treating any content the agent reads from external/untrusted sources as data rather than instructions, using permission scoping so the agent can't take actions beyond what a task genuinely requires, and logging/monitoring agent actions for review. No single technique fully eliminates the risk — it's an active area of ongoing security research.
Why might a model's context window size not be the only limiting factor for how well it uses long context?
Even within a technically supported context window, models can show uneven attention to different parts of a long input — sometimes described informally as being better at using information near the beginning or end of the context than content buried in the middle. This means simply fitting more text into the context window doesn't guarantee the model will weigh all of it equally well, which is a practical consideration when designing prompts or RAG systems with long context.
How would you evaluate whether a prompt change actually improved output quality, rather than just seeming better on a few examples?
Rather than spot-checking a handful of examples, build a small evaluation set of representative test prompts with defined expected qualities (accuracy, format compliance, tone), run both the old and new prompt against that set, and compare results systematically — ideally with some cases scored by a human rather than only by the model itself, since a model judging its own output can share the same blind spots as the generation step.