Zero-Shot vs Few-Shot Prompting: When Examples Help or Hurt
Zero-shot prompts give only instructions; few-shot prompts add examples. Learn when examples improve output, and when they cause copying, bias and drift.
In this article
Zero-shot prompting means giving a model only instructions and the input; few-shot prompting adds a small number of worked examples so the model can copy the pattern. Use zero-shot when the task is easy to describe in words, and add examples when the format, style or decision boundary is easier to show than to explain.
The terms come from the GPT-3 paper (Brown et al., 2020, "Language Models are Few-Shot Learners"), which showed that a large model could pick up a task from a few demonstrations in the prompt, without any retraining. That ability is usually called in-context learning. Six years on, models follow plain instructions far better than GPT-3 did, so the question is no longer "do I need examples to make this work at all" but "do examples make this better or worse for my case".
Key takeaways
- Start zero-shot with clear instructions. Add examples when outputs are inconsistent in format or style.
- Examples are copied more literally than most people expect: length, tone, structure, even topic.
- Vary your examples, balance the labels, and do not always put the same class last.
- Examples complement instructions; they rarely replace them.
What does a zero-shot prompt look like?
Extract the company name and the job title from this email signature.
Return them as "Company: ..., Title: ...". If either is missing, write "unknown".
Signature:
Maria Lopez
Head of Data Platform | Northwind Logistics
+1 555 0100
A modern model gets this right almost every time. The task is common, the output format is described exactly, and there is a rule for the missing case. Adding examples here would cost tokens and buy nothing.
What does a few-shot prompt look like?
Now a task where words fall short. You want product descriptions rewritten in your brand's voice, which is dry, slightly funny and allergic to adjectives. You could try to describe that voice. Or you could show it:
Rewrite product descriptions in our house style.
Original: "This premium, ultra-durable water bottle keeps your drinks
ice-cold for an amazing 24 hours!"
Rewrite: "Keeps water cold for 24 hours. Survives being dropped
down stairs, which we tested more than we needed to."
Original: "Experience unparalleled comfort with our luxurious,
ergonomically designed office chair."
Rewrite: "An office chair that supports your lower back. You will
stop noticing it, which is the point."
Original: "{{new_description}}"
Rewrite:
Two examples communicate more about the voice than a paragraph of adjectives ("witty but understated") ever would. This is the strongest case for few-shot: style, tone and format that are easy to recognize and hard to specify.
When do examples clearly help?
From experience, examples pay off in a few recurring situations:
Unusual output formats. A custom markup, a specific way of writing dates, a citation style your team invented. One example of the finished output removes a lot of guesswork.
Fuzzy classification boundaries. If "complaint" vs "feature request" depends on your team's conventions, show a borderline case labeled the way you want it. Pair it with a written rule, because the model may not generalize from one instance.
Style transfer. As above. Voice is the hardest thing to describe and the easiest to demonstrate.
Smaller or older models. The less capable the model, the more it leans on demonstrations. If you run a small open-weight model locally, few-shot often matters a lot more than it does with a frontier model.
When do few-shot examples hurt?
This is the part most tutorials skip. Examples are strong signals, and strong signals bring side effects.
The model copies the wrong things
Suppose both of your examples happen to be about kitchen products and around 30 words long. Your outputs will drift toward 30 words, and you may see kitchen metaphors creeping into descriptions of laptop bags. The model cannot tell which features of the example are the point and which are accidents. It copies all of them a bit.
The fix is variety. Make examples differ in length, topic and structure along every dimension you do not care about, and keep them consistent only on the dimension you do care about.
Label imbalance and ordering bias
Research on few-shot classification has documented this directly. Zhao et al. (2021, "Calibrate Before Use") found that GPT-3's predictions were skewed toward labels that appeared more often in the prompt and toward the label of the last example. Newer models are less fragile, but the tendency has not disappeared.
So if your sentiment prompt has three positive examples and one negative, do not be surprised when neutral messages come back positive. Balance the classes and shuffle the order, or at least make sure the last example is not always the same label.
Examples that contradict the instructions
A surprisingly common bug:
Summarize the article in at most two sentences.
Example article: ...
Example summary: The company reported higher revenue this quarter.
It also announced a new CEO. Analysts expect the stock to rise.
The instruction says two sentences, the example has three. When instructions and examples disagree, models often follow the example. Audit your examples against your own rules before blaming the model.
Examples that lock in a narrow pattern
If every example of "extract action items" shows exactly two items, the model may start producing two items even for a meeting that had five. Include an example with zero items and one with many. Edge cases in your examples teach the model that edge cases are allowed.
Do the labels in examples even need to be correct?
Interesting question, and the research answer is "less than you would think, but do not rely on that". Min et al. (2022, "Rethinking the Role of Demonstrations") found that, for a range of classification tasks, randomly swapping the labels in demonstrations hurt performance only modestly; much of the benefit came from showing the input distribution, the label space and the format. That is a useful insight into how examples work. It is not permission to use sloppy examples in production, where a wrong example can teach a wrong rule on exactly the borderline cases you care about.
How many examples should you use?
There is no universal number. In practice:
- One example shows a format, but invites heavy copying.
- Two to five well-chosen, varied examples handle most tasks.
- Beyond that, gains usually flatten, and each example adds cost and latency on every call.
The only reliable way to choose is to test. Keep a small evaluation set (twenty or thirty real inputs with known good outputs), try zero, two and five examples, and look at both accuracy and the kinds of errors. Sometimes more examples fix one failure mode and introduce another.
Should examples replace instructions?
Rarely. The strongest prompts pair a short, explicit rule with examples that illustrate it. Instructions tell the model what matters; examples show what it looks like. If you rely on examples alone, the model has to reverse-engineer your intent, and it will sometimes infer a different rule than the one you had in mind.
A practical pattern is to mark examples clearly so the model does not confuse them with the real input:
Classify each support ticket as "bug", "billing" or "question".
A ticket that reports an error message counts as "bug" even if the
user phrases it as a question.
<examples>
<example>
Ticket: Why do I get "Error 502" when I export a report?
Label: bug
</example>
<example>
Ticket: Can I pay annually instead of monthly?
Label: billing
</example>
<example>
Ticket: Is there a keyboard shortcut for search?
Label: question
</example>
</examples>
Ticket: {{ticket}}
Label:
Notice the first example is the borderline case the written rule describes. That is deliberate: the instruction states the rule, the example proves you meant it.
A quick decision guide
Ask yourself whether a competent new colleague could do the task from your written instructions alone. If yes, go zero-shot and spend your effort on clear instructions and format. If they would ask "can you show me one?", add examples, vary them, balance them, and check they agree with your rules.
For more on this, the few-shot topic breaks the decisions down further, and the companion piece on chain-of-thought prompting covers what happens when the examples include reasoning steps, not just answers. If you are new to the basics, start with what prompt engineering is.
Frequently asked questions
What is the difference between zero-shot and few-shot prompting?
A zero-shot prompt describes the task with instructions only. A few-shot prompt also includes a handful of worked input/output examples so the model can infer the pattern from them.
How many examples should a few-shot prompt have?
Usually two to five well-chosen examples are enough. Add more only if testing shows a gain, and make sure they cover different cases rather than repeating the same one.
Can few-shot examples make results worse?
Yes. Models tend to copy surface features of examples, such as length, phrasing and topic, and can be biased toward labels that appear more often or last in the example list.
What is one-shot prompting?
One-shot prompting is few-shot prompting with a single example. It is useful for showing a format, but one example makes it easy for the model to copy that example too closely.
- #few-shot prompting
- #zero-shot prompting
- #in-context learning
- #examples