Few-shot prompting is the practice of including a small number of worked examples directly in a prompt, so the model can infer the pattern you want before it generates its output. It's one of the oldest tricks in the large language model playbook, and it still works remarkably well. It also fails in predictable ways that most teams only discover after they've shipped something to production and started paying for it.
The core idea is simple. Instead of writing a detailed instruction, you show the model two or three examples of input and the corresponding output you expect. The model extracts the pattern and applies it to your actual query. No fine-tuning required, no extra training data, no model deployment. Just tokens.
Why few-shot prompting works at all
Large language models are trained on so much text that they've seen almost every format, style, and task type imaginable. When you supply examples, you're not teaching the model anything new. You're narrowing the probability distribution of what it outputs. The model already knows how to write a JSON object, classify a sentiment, or summarise a document. The examples tell it which version of that knowledge you want right now.
This is why few-shot prompting is most effective on tasks with a clear, consistent output format. Extracting structured data from unstructured text is a classic case. Classifying customer support tickets into a fixed taxonomy is another. Transforming one writing style into another, such as converting internal notes into formal client-facing summaries, responds well to examples because the delta between input and output is learnable from 3 to 5 demonstrations.
The quality of the examples matters far more than the quantity. A single well-constructed example that covers the edge cases your task actually produces will outperform five sloppy ones that represent only the easy cases. Teams that pick examples from the easiest instances in their dataset often find the model performs well on those and badly on everything else.
When few-shot prompting genuinely earns its token cost
Few-shot prompting pays off when three conditions hold at once. First, the task has a consistent, repeatable output structure. Second, zero-shot or instruction-only prompting produces outputs that are close but not quite right, with consistent errors that examples could correct. Third, the context window cost of including examples is small relative to the value of the improved output.
Structured extraction is the clearest win. If you're pulling product attributes from unstructured supplier descriptions, a few examples demonstrating how to handle missing fields, inconsistent units, and ambiguous values will save you significant post-processing time. The model learns your tolerance for ambiguity from the examples, not from a paragraph of rules.
Format enforcement is another strong use case. If your downstream system expects a specific JSON schema, showing the model two correctly formatted outputs is more reliable than describing the schema in prose. Prose descriptions of schemas introduce ambiguity; examples are unambiguous.
Tone and register transfer also responds well to examples. If your brand voice has specific characteristics that are hard to articulate but obvious in practice, three examples of input text and the desired rewritten output can calibrate the model faster than any style guide. This is relevant for Australian enterprises localising global content, where the register gap between US English and Australian business English is real but subtle.
When it wastes tokens and adds noise
Few-shot prompting fails, or at least underperforms, in several well-documented situations. Open-ended reasoning tasks are the most common trap. If you're asking a model to analyse a business problem, generate creative options, or produce a nuanced recommendation, providing examples can actually constrain the output in ways you don't want. The model anchors to the structure and length of your examples rather than reasoning freely. You get answers that look like your examples rather than answers that are genuinely good.
Tasks where the correct output varies significantly across instances are another failure mode. If your examples are drawn from a narrow slice of the input space, the model will generalise from them incorrectly. It's the same overfitting problem that plagues supervised machine learning, applied to in-context learning. The model doesn't know your examples are unrepresentative.
Long examples in long context windows compound AI token costs quickly. If each of your four examples is 800 tokens, you've spent 3,200 tokens before the actual query appears. For high-volume inference pipelines, that cost accumulates fast. The break-even point depends on your model provider's pricing and the frequency of the task, but the calculation is worth doing before you lock in a prompt template.
Label bias is a subtler problem. Research into few-shot prompting has repeatedly shown that models are sensitive to the distribution of labels across examples. If three of your five classification examples are positive instances and only two are negative, the model will classify ambiguous cases as positive at a higher rate. Balancing your examples by output class is a basic hygiene step most teams skip.
Practical guidance for Australian enterprise teams
Start with zero-shot. Before you add examples, try a well-written instruction-only prompt. Many tasks that seem to need examples actually just need a cleaner instruction. Adding examples to a prompt that already works well rarely improves output quality and always increases cost.
When you do add examples, choose them deliberately. Cover the edge cases your task actually encounters, not the easy cases that already work. Include at least one example where the correct output involves a judgment call or an ambiguity, and show how you want the model to handle it.
Keep examples short. The ideal few-shot example demonstrates the pattern without being exhaustive. If your examples are running to several hundred tokens each, consider whether you're trying to demonstrate too much in each one, or whether the task is complex enough to warrant fine-tuning rather than prompting.
Test example order. Models are sensitive to recency. The last example before the actual query has disproportionate influence. Put your most representative example last. This is a simple change that measurably improves consistency on classification and extraction tasks.
Version your prompt templates. Few-shot prompts are code. They should be stored in source control, reviewed when examples are added or changed, and tested against a fixed evaluation set before deployment. Teams that treat prompt templates as ad hoc text files discover their model behaviour has changed silently when someone edited an example to fix one case and broke three others. The same discipline that applies to AI model versioning applies to the prompts that drive those models.
The interaction with model size
Smaller models benefit more from few-shot examples than larger ones do. A large frontier model like GPT-4o or Claude 3.5 Sonnet can often infer your intent from a precise zero-shot instruction. A smaller, faster model running on-premises or at the edge may need examples to compensate for its narrower training. This has real implications for Australian teams running local inference for latency or data-sovereignty reasons: the cost of few-shot prompting in tokens may be justified precisely because the model is smaller and cheaper per token.
The practical upshot is that your prompting strategy and your model choice interact. Don't optimise your few-shot examples for one model and then switch models without retesting. The examples that work well for a 7B parameter model may actively hurt performance on a frontier model, and vice versa.
Few-shot prompting is one of the highest-leverage controls available to teams working with off-the-shelf language models. It's also easy to over-engineer. The teams that use it well treat it as a precision instrument, applied to specific failure modes, tested against real data, and governed like any other production asset.

