Skip to content

LLMs in 2026 ​

Module objectives

  • Know how to situate a model: its size, its architecture, its context window, its reasoning level, its ability to call tools
  • Understand just enough of how an LLM works to choose the right model later, depending on the role you assign it

This module is part of the prerequisites, and we treat it as such. The topic is covered in depth elsewhere, often better than we could do here. Our goal is not to compete with those resources, but to bring everyone to the same level before tackling the reconstruction. We therefore deliberately stay at the surface, and point to reference reading for those who want to go deeper, starting with Andrej Karpathy's introduction to large language models, a one-hour warm-up on the question "what is an LLM", extended by his detailed 2025 course for anyone who wants the full training stack.

What a model does, fundamentally ​

A language model predicts the next word. More precisely, it predicts the next token, that is, the next fragment of text, based on everything that precedes it. The text you give it is first split into tokens, then the model produces one token at a time, each one added to the input to predict the next. The ability to answer a question, write code, or reason emerges from this simple mechanism applied at very large scale.

This way of working explains two things that will serve us throughout the training. On the one hand, the model only "knows" what is in its input or what it learned during training. On the other hand, the quality of what you put in bears directly on the quality of what comes out. That is where "prompt engineering" really matters.

The ground has nonetheless shifted since the early guides. The classic techniques (in-context examples, chain of thought, self-consistency) remain a useful foundation, but their weight has changed: reasoning models now produce the chain of thought on their own, which makes manual guidance less necessary; tool calling has become a native feature rather than a phrasing trick; and attention has shifted from optimizing an isolated prompt to organizing an agent's entire context, what is now called "context engineering". It is a thread that leads straight to building agents, and that we will pick up again in act 2.

The context window ​

The model cannot take an infinitely long text into account. It has a context window, a maximum number of tokens it can consider at once. This window has grown considerably in recent years, to the point of exceeding a million tokens on some recent models.

Yet this growth does not solve the problem, because models make poor use of information located in the middle of a long context, a phenomenon known as lost in the middle. Filling the window is therefore not enough: what matters is what you put in it, and where. This observation alone drives much of the work on context that we will carry out in Act 2.

A nuance about recent models

The lost in the middle effect no longer fully holds on the latest models. On needle-in-a-haystack retrieval tests, Anthropic's recent models achieve near-perfect recall, above 99% as early as Claude 3 Opus, regardless of where the information sits in the context. The effect therefore weakens considerably for simple fact retrieval; it remains more pronounced as soon as the task requires reasoning over several pieces of information scattered across the context. The practical lesson does not change: being careful about what you put in the window, and where, still pays off.

Mixture of experts ​

Many recent models rely on an architecture called mixture of experts (Mixture of Experts, or MoE), for which Hugging Face offers an illustrated overview. The idea is not to activate the whole network at each token, but only a small part, chosen dynamically. A model can therefore advertise a very high total parameter count while activating only a fraction of them at each step.

The practical consequence is that you need to distinguish total parameters from active parameters. The former provide information about the model's capacity and the memory needed to load it; the latter about its compute cost and speed. Two models announced with the same number of parameters can behave very differently depending on this distinction.

A few examples from open models, where the gap between total and active parameters is obvious whenever a MoE is involved:

YearModelArchitectureTotal parametersActive parameters
2025Kimi K2 (Moonshot AI)MoE1T32B
2024DeepSeek-V3MoE671B37B
2025Llama 4 Maverick (Meta)MoE400B17B
2025Qwen3-235B-A22BMoE235B22B
2026Gemma 4 26B A4B (Google)MoE26B4B
2026Gemma 4 31B (Google)Dense31B31B
2025Qwen3-32BDense32B32B
2025Mistral Small 3Dense24B24B

On a dense model, the two columns are identical: the whole network is activated at each token. On a MoE, the gap can be considerable: DeepSeek-V3 loads 671 billion parameters but only activates 37 at each step. Large proprietary models (GPT, Claude, Gemini) are widely assumed to rely on MoE too, but their architecture is not disclosed, so here we stick to open models.

The model name often gives a first hint. The A<n>B suffix, for Active <n> Billion, announces the number of active parameters: “Gemma 4 26B A4B” refers to 26 billion parameters in total but 4 billion active, and “Qwen3-235B-A22B” 235 billion for 22 active. A dense model never carries this suffix, since active and total parameters are one and the same. Note, however, that this convention is not universal: Kimi K2 or DeepSeek-V3 are indeed MoE without displaying it in their name. The reliable reflex is to check the model card, where both counts are stated.

The reasoning level ​

A model can answer on the spot, or take the time to “think” before concluding. Since late 2024, a family of reasoning models has made this second way a mode of its own: before producing its answer, the model generates a long chain of intermediate tokens, a step-by-step reflection that is not necessarily shown to the user, but that clearly improves results on difficult tasks (mathematics, code, planning, multi-step problems).

This is a fundamental change. Until then, you improved a model mainly by training it longer on more data. Here, you gain quality by letting it spend more computation at answer time, what is called test-time compute. The movement unfolded quickly: OpenAI o1 in September 2024, DeepSeek-R1 in January 2025, then the extended thinking of Claude 3.7 Sonnet in February 2025. Within a few months, inference-time reasoning became a standard.

This extra thinking has a cost: it consumes many tokens and lengthens response time. That is why most of these models let you adjust the reasoning effort, from a fast and economical mode up to deep thinking. The whole art is to use it only when it adds something. This is directly linked to choosing the model by role, below: a planner benefits from reasoning at length, a constrained executor does not need it and would be needlessly expensive.

Tool calling ​

A model that only produces text cannot act. For it to become an agent, it must be able to trigger actions: read a file, run a command, query an API. This is the role of tool calling. The model does not perform the action itself; it produces a structured request, which the harness executes, before sending the result back to it.

This capability is recent by the standards of LLM history. It was first explored on the research side, with the ReAct paradigm (reason then act) in late 2022, then Toolformer in early 2023, before becoming a full-fledged API feature: OpenAI introduced function calling in June 2023, and Anthropic launched tool calling on Claude in beta in late 2023, before its general availability in May 2024. In less than two years, we thus went from a simple text model to an agent capable of acting.

This capability is the prerequisite for everything that follows. A harness is precisely what organizes this loop between the model and the tools, and a good part of the training consists in rebuilding its inner workings.

Where to find models and their specifics? ​

All the figures in the previous table, and many more, can be read in the same place. Open models are now published on the Hugging Face Hub, a platform that hosts the model weights, their documentation, and a way to try them out. It's the first reflex when you're trying to situate a model.

Each model has a model card, a README written by its publisher. It holds what matters for this module: the model size, its architecture (dense or MoE, number of experts), its context window, the supported languages and modalities, results on the major benchmarks, and the usage license.

For the technical details the model card sometimes omits, the model's config.json file gives the raw configuration: internal dimensions, number of layers, number of experts and number of active experts for an MoE. That's where you confirm, numbers in hand, the gap between total and active parameters mentioned earlier.

Finally, the Hub is not just for browsing. Its filters let you explore models by task, size, or license; comparative leaderboards help you find your way in a fast-moving catalog; and quantized versions (often in GGUF format), being lighter, make some models runnable on a modest machine.

Choosing a model according to its role ​

All of this has a practical aim. When we build agents, we will assign them distinct roles, and these roles do not call for the same model. An agent tasked with planning benefits from relying on a solid model, capable of reasoning at length. An agent that executes a repetitive, well-scoped task benefits, for its part, from a fast and economical model.

The best way to build an intuition is still to try things out. Free websites let you submit the same prompt to two models and compare their answers side by side. The most useful is LMArena (formerly Chatbot Arena): its side-by-side mode lets you choose the two models to compare, without even creating an account; its battle mode, where two anonymous models respond and you vote, also feeds a comparative leaderboard. Hugging Face Chat and OpenRouter offer the same kind of testing across a large catalog. Nothing beats running your own prompt on a fast model and on a reasoning model to feel, concretely, what each brings and what it costs. This is precisely what we will see in the first part of the training.

References ​

Tools ​

  • LMArena: compare two models side by side on the same prompt.
  • Hugging Face, Chat: try open models online.
  • OpenRouter: access a wide catalog of models through a single interface.
© 2026 Loic Gouarin, Max Beligné. This content is published under the CC BY-SA 4.0 license.