LLMOps, In Depth · Part 3 of 6

Open Models: MoE, Reasoning & Fine-tuning

Opening the box — the architecture inside, how a model is trained to think, the math of shrinking it to fit your hardware, and how to specialize it on your own GPU.

By PrithvirajPart 3 of 6~38 min read

In the first two posts the model was a service you called over the internet. You sent text in and got text back, and you never saw what was inside. This post is about the other option: models where the company publishes the actual numbers that make up the model — its weights — so you can download them and run them yourself. These are called open-weight models. Once you have the weights, the model stops being something you rent and becomes something you own: you can look inside it, shrink it, retrain part of it, and run it entirely on your own machines.

There are four things you must understand to work with these models, and this post covers them one at a time: the design trick that lets a huge model run cheaply (Mixture-of-Experts), the training that makes a model work through a problem step by step (reasoning), the math that shrinks a model so it fits on your hardware (quantization), and the technique that lets you adapt a model to your own task (fine-tuning).

The Deep Learning series' Part 3 explained the transformer and how models are trained from human feedback. Here we look at the same machinery from the point of view of the person who has to run it in production — what each choice does to cost, privacy, speed, and how specialized the model can be.

Why owning the weights changes your decisions
Calling a hosted model over an API is easy: no servers to manage. But you pay for every word processed, your data leaves your building, and the only way to change the model's behavior is to reword your prompt. Running open weights flips all three: you pay a fixed price for hardware instead of a per-word fee, your data never leaves your control (so you can meet rules like HIPAA for health data or GDPR for EU privacy), and you are free to retrain the model. The cost of that freedom is that you now have to run the model yourself (Part 4), test it yourself (Part 2), and keep it safe yourself (Part 6). This post is about how to use that freedom well.

Part 1 · Open Models

The open-weight landscape

"Open" is not one thing — it is a range. At one end, a project releases everything: the weights, the training data, and the training code, under a license that lets anyone use it freely. At the other end, a company releases only the weights, under its own custom license with rules attached. That second, more common case is what people mean by "open weights." When you pick a model to run, four practical questions matter: does the license let me use it in a commercial product, what design does it use (explained below), how much text can it read at once (its context window), and what is it especially good at.

Two design words appear in the table. A dense model runs every one of its weights on every word — simple and predictable. A Mixture-of-Experts model (MoE for short, explained fully in Part 2) stores many weights but uses only a small slice of them on each word, so it is cheaper to run than its size suggests.

FamilyNotable formDesignLicense character
Llama (Meta)3.1 dense (8B/70B/405B); Llama 4 uses MoE (Scout/Maverick) with a very long context windowDense + MoECustom community license (caps monthly active users)
Qwen (Alibaba)Qwen3 in both dense and MoE forms (e.g. 235B total / ~22B used per word), can switch thinking on and offDense + MoEMostly Apache-2.0 (very permissive)
DeepSeekV3: 671B total / ~37B used per word; R1 adds reasoning on top of the V3 baseFine-grained MoEMIT (very permissive, for the weights)
MistralMistral 7B (dense); Mixtral 8×7B and 8×22B (MoE)Dense + MoEApache-2.0 (older) / research-only (some newer)
Gemma (Google)Gemma 3 (1B/4B/12B/27B), handles images too, 128K contextDenseCustom Gemma terms
GLM (Zhipu)GLM-4.x, large MoE, aimed at reasoning and tool-using tasksMoEMIT (for the weights)
Operator's note — the license is a shipping decision, not fine print
Apache-2.0 and MIT are very permissive licenses: they let you build and sell products with almost no legal friction, which is why Qwen, DeepSeek, and Mistral-7B/Mixtral are safe defaults for commercial work. Llama's community license adds a rule that kicks in above a certain number of monthly users; Gemma has its own terms. Read the actual license against what you plan to ship — it limits your product just as firmly as any hardware limit. (The model details here reflect public information as of early 2026 and change fast; verify against current model cards.)

1.1 Why models end up good at different things

One family is best at writing code, another at chatting in many languages, another at using tools on your behalf. These strengths are not accidents. They are decided at three stages, all built on top of the same basic transformer:

  • What it reads during first training (the pre-training data mix): a coding model is fed trillions of words of source code and technical docs; a generalist is fed a balanced diet of web pages, books, and text in many languages.
  • How it handles long inputs: models that can read very long documents use tricks in how they track word positions (named RoPE scaling and YaRN) to avoid the known problem where a model ignores facts stuck in the middle of a long input.
  • What it practices after first training (post-training): a tool-using model is trained on recordings of real tool calls — the JSON requests, the API calls, the results — while a chat model is trained on examples of people ranking which reply is better.

Part 2 · Open Models

Mixture-of-Experts, in full

Almost every very large open model uses one trick to stay affordable to run. The trick answers a question that sounds impossible: how do you get the knowledge of a 671-billion-weight model while paying to run only about 37 billion weights per word? The trick is called Mixture-of-Experts, or MoE ("mixture of experts" — a set of specialist sub-networks the model chooses between).

2.1 What an "expert" really is

Recall the transformer block from the Transformer post: first an attention step that lets each word look at the others, then a feed-forward network — a small stack of math that transforms each word on its own. That feed-forward network is usually written FFN (feed-forward network). In a normal dense model there is one big FFN in each block, and every word passes through all of it. In an MoE model, that single FFN is replaced by many smaller FFNs sitting side by side. Each one of these smaller FFNs is called an expert. So "expert" is not something clever — it is just one FFN out of many.

A small extra piece decides which experts each word should use. It is a tiny layer that scores every expert and picks the best few. It is called the router. The attention step is left shared and untouched — it keeps the flow of context between words — and only the FFN part is split into experts.

a word \(x\) finishes the attention step →
Router \(W_g\) — scores all experts, keeps the top \(k\)
Expert 1
on
Expert 2
Expert 3
Expert 4
on
…
Expert N
combine only the experts that are on → pass to the next layer
The plain idea, in one line
The model keeps a very large set of experts on hand (that is its total weight count), but for each word it turns on only a couple of them (that is its active weight count). It pays to run only the few it turns on, yet it still holds all the knowledge stored in the full set. That is how DeepSeek-V3 can hold 671 billion weights of knowledge while costing about 37 billion weights of work per word.

2.2 The routing math

Here is what actually happens for one word. The word arrives as a list of numbers, written \(x\). The router multiplies it by its weight matrix \(W_g\) to get one score per expert. It keeps the highest \(k\) scores (the "top-k"), turns those scores into fractions that add up to 1 using softmax, and produces the final output as those fractions times the outputs of just the chosen experts:

$$h(x)=x W_g,\qquad G(x)=\text{softmax}\big(\text{TopK}(h(x),k)\big),\qquad y=\sum_{i\in\text{TopK}} G(x)_i \cdot E_i(x)$$

Each expert \(E_i\) is just an ordinary FFN, \(E_i(x)=\text{Activation}(xW_{i,1})W_{i,2}\). "Top-k" means keep the k best-scoring experts; the softmax turns their raw scores into weights that sum to 1 so the combination is a proper weighted average. For example, Mixtral has \(N=8\) experts and picks \(k=2\) of them per word: about 47B weights in total, but only about 13B doing work on any given word.

Worked example — active vs total, and where the saving comes from

Say an MoE layer has \(N=8\) experts. Each expert is an FFN with \(P_e\) weights, and the router picks \(k=2\) of them per word.

$$\text{total FFN weights}=8P_e,\qquad \text{weights actually run per word}=2P_e$$

The word is only multiplied against \(2P_e\) weights instead of all \(8P_e\) — that is a 4× saving in computation (8 divided by 2). Meanwhile the model still stores all \(8P_e\) weights, so it keeps all that specialized knowledge available. If you use more, smaller experts (DeepSeek does exactly this), the ratio of knowledge-stored to work-done climbs even higher. But note one catch: memory (the VRAM on your GPU) is set by the total weight count, because you must load every expert into memory even though you run only a few. So MoE buys cheap computation at the price of expensive memory — a fact that shapes serving in Part 4.

2.3 The problem MoE creates: keeping the experts balanced

If you just let the router learn freely, it cheats. It discovers a few experts it likes and sends almost every word to them, so the rest of the experts never get used or trained. Those unused experts are called dead experts. The general problem of keeping traffic spread evenly across experts is called load balancing. Two fixes are standard:

  • A balancing penalty added to training (the auxiliary load-balancing loss, from the Switch Transformer work): an extra term in the training goal that punishes lopsided expert usage across a batch of words, pushing the router to share the load.
  • Balancing without a penalty term (the auxiliary-loss-free method, from DeepSeek-V3): instead of a penalty, it nudges a small per-expert offset up or down to even out traffic, so the main training goal is not distorted — a newer, cleaner approach.

DeepSeek adds two more ideas. Fine-grained experts means using many small experts instead of a few big ones, so each can specialize more narrowly. Shared experts means keeping one or two experts always turned on for every word, to hold the common knowledge every word needs, which frees the routed experts to specialize harder.

Operator's implication — MoE reshapes your hardware bill
Because the active weight count drives speed but the total weight count drives memory, an MoE model runs fast yet is hungry for memory to host. A 671B-total / 37B-active model produces words at roughly 37B speed but still needs enough VRAM to hold all 671B weights (usually after shrinking them, see next). This is why big MoE models are spread across several GPUs (Part 4) and why quantization — coming up next — is often required rather than optional for them.

Part 3 · Open Models

Reasoning models — training a model to think first

A plain language model just predicts the next word that is most likely to come — a fast, one-shot guess. For hard logic problems that quick guess is often wrong. A reasoning model is trained to write out a long train of thought before it gives its final answer, checking and correcting itself as it goes. It moves effort from training time to answer time — the time when it is actually producing your reply, which people call test time.

3.1 Thinking out loud, as a trained habit

The model writes its working-out inside special marker tokens, <think>...</think>, and then writes the final answer after. Why does writing out its thoughts help? Because of a fact from the foundations: each word the model generates becomes part of the input for the next word. So when it writes "step 1… step 2… wait, that is wrong, let me redo it…" it is literally giving itself more chances to compute and more notes to build on before committing to an answer. This is called chain-of-thought — the model reasoning in a visible chain of steps. More thinking words means more work done at answer time, which means better accuracy on math, code, and logic — up to a point, after which extra thinking stops helping.

3.2 How the thinking is trained: rewards for getting it right

DeepSeek-R1 showed you can teach a model to reason mostly by trial and error with rewards, a method called reinforcement learning. Instead of copying human-written reasoning, the model tries to solve problems whose answers can be checked automatically — a math problem with a known result, or code that must pass a test. When it solves one, it gets a reward. Over many tries it discovers, on its own, that checking its steps and going back to fix mistakes gets more rewards.

Engineering track — GRPO in plain terms
The specific reinforcement-learning method behind R1 is GRPO, which stands for Group Relative Policy Optimization. An older method called PPO needs a second helper network (a "critic") to judge how good each partial answer is — and that helper is expensive to train. GRPO throws the critic away. For each question it generates a group of \(G\) different answers, scores every one with a simple rule-based reward \(r_i\) (Is the answer correct? Is it formatted properly?), and then judges each answer against the group's own average. The formula for how much better than average answer \(i\) is: \(A_i=\frac{r_i-\text{mean}(r)}{\text{std}(r)}\), where \(\text{mean}(r)\) is the group's average score and \(\text{std}(r)\) is how spread out the scores are. Answers that beat their group's average get reinforced; answers below average get pushed down. No critic network, cheaper training, and the reward is objective — a real checker, not a guessed opinion.
Tiny worked example — how GRPO scores a group

Suppose for one math question the model generates \(G=4\) answers and the checker gives them rewards \(r = [1, 0, 1, 0]\) (1 = correct, 0 = wrong).

$$\text{mean}(r)=\frac{1+0+1+0}{4}=0.5,\qquad \text{std}(r)=0.5$$

The advantage of a correct answer is \(A=\frac{1-0.5}{0.5}=+1.0\); the advantage of a wrong one is \(A=\frac{0-0.5}{0.5}=-1.0\). Training then makes the two correct answers more likely and the two wrong answers less likely. No separate judge network was needed — the group graded itself.

3.3 Hybrid reasoning — one model with a thinking dial

Early on, a reasoning model was a separate download from the normal model. Newer models (Qwen3, Claude, and others) combine both in one and let you set how much thinking you want. This combined kind is called a hybrid reasoning model, and the setting is a thinking budget — a cap on how many thinking words it may use. Set the budget to zero and it answers instantly like a normal model; raise the budget and it works through the problem first.

budget = 0
[input] → [answer]  — fast and cheap, just like a normal model
vs
budget > 0
[input] → [<think> reason, check, go back and fix </think>] → [answer]
Operator's implication — thinking is a cost and speed dial
Thinking words cost money (you are billed for them) and add delay, often making a reply take 3 to 10 times longer. Treat the thinking budget as a per-request setting: turn it up for a hard coding or math task, set it to zero for a simple sorting or fact-pulling task. Sending easy requests to zero-thinking (or to a smaller model) and only hard ones to full thinking is one of the largest cost savings available to you — something we put real numbers on in Part 6.

Part 4 · Open Models

Quantization & precision — making a model fit

Every weight in a model is stored as a number with a certain number of digits of accuracy. The more digits, the more memory it takes and the slower it is to move around. Storing each weight using fewer bits — trading a little accuracy for a lot of memory savings — is called quantization. It is what lets a model that supposedly "needs" a data-center GPU run on a laptop.

4.1 The memory math

The memory a model's weights take is simply the number of weights times how many bytes each weight uses:

$$\text{memory} \approx (\text{number of weights})\times(\text{bytes per weight})$$

The precision level tells you the bytes per weight. Some names you will see: FP16 and BF16 are two 16-bit (2-byte) formats — "full" precision for running a model. INT8 stores each weight as an 8-bit (1-byte) whole number. INT4 and NF4 use just 4 bits (half a byte) each. Here is the same 8-billion (8B) and 70-billion (70B) weight model at each level:

PrecisionBits per weight8B model weights70B model weights
FP16 / BF1616≈ 16 GB≈ 140 GB
INT88≈ 8 GB≈ 70 GB
INT4 / NF44≈ 4 GB≈ 35 GB

Check the math: 8 billion weights × 2 bytes = 16 GB at FP16; at 4 bits (half a byte) it is 8 billion × 0.5 = 4 GB. Two things get cheaper at once. First, capacity: the model fits in less VRAM. Second, speed: producing words is limited mostly by how fast weights can be moved from memory into the compute cores, and 4-bit numbers move four times faster than 16-bit ones. So a 4-bit model is about a quarter of the size and noticeably faster, at a small cost to quality.

Engineering track — how the shrinking actually works
A smooth 16-bit weight \(x\) that lives somewhere between a minimum and maximum value gets snapped to one of a limited set of whole-number bins: $$q=\text{round}\!\left(\frac{x}{S}\right)+Z$$ Here \(S\) is a scale (how wide each bin is) and \(Z\) is a zero-point (which bin represents zero). When the weight is needed for computation, this whole number is converted back toward a float. The whole skill is choosing \(S\) and \(Z\) — and deciding which important weights to keep at higher accuracy — so that the rounding barely changes the model's output. That is exactly what the different methods below are competing to do best.

4.2 The formats you will actually run into

Each name below is a different way of storing or producing quantized weights. Here is a one-line plain description of each:

FormatWhat it isWhere it is used
GGUFThe quantized file format used by the llama.cpp project (offers levels from Q2 to Q8)Running models locally on a CPU, a Mac, or a consumer GPU (via Ollama, LM Studio)
GPTQA method that shrinks an already-trained model to 3–4 bits in one pass, using extra math about how sensitive each weight isServing quantized weights on a GPU
AWQ"Activation-aware" — it looks at which weights matter most in practice and keeps those more accurateHigh-quality 4-bit inference on a GPU
NF4 (from the bitsandbytes library)A 4-bit format ("NormalFloat") whose bins are placed to fit the bell-curve shape that weights naturally followQLoRA fine-tuning (next section)
FP8An 8-bit floating-point format (variants E4M3 and E5M2) that recent GPUs support directly in hardwareHigh-throughput serving; DeepSeek-V3 even did its training in FP8

Part 5 · Open Models

Fine-tuning: LoRA & QLoRA on your own GPU

Adjusting a base model so it fits your domain, your tone, or your specific task is called fine-tuning. The thorough version — updating every weight in the model — is called full fine-tuning, and it needs a cluster of data-center GPUs, so most teams never do it. The technique that made fine-tuning affordable for everyone is LoRA, and a further trick called QLoRA lets you fine-tune a capable model on a single gaming-grade GPU.

5.1 LoRA — train a tiny add-on, freeze everything else

Instead of changing a model's giant weight matrix \(W\) directly, LoRA (which stands for Low-Rank Adaptation — "low-rank" is explained in a moment) leaves \(W\) frozen and learns a small correction to add on top. It writes that correction \(\Delta W\) as the product of two thin matrices, \(B\) and \(A\):

$$W' = W + \Delta W = W + \frac{\alpha}{r}\,B A,\qquad B\in\mathbb{R}^{d\times r},\ A\in\mathbb{R}^{r\times k},\ r\ll d,k$$

Only \(A\) and \(B\) are trained; the huge \(W\) never moves. The number \(r\) is the rank — how thin the two matrices are — and "low-rank" just means \(r\) is chosen much smaller than the matrix's normal dimensions. Concretely: take a \(4096\times4096\) weight matrix, which has \(4096\times4096\approx16.7\) million weights. With rank \(r=16\) you instead train \(B\) (size \(4096\times16\)) plus \(A\) (size \(16\times4096\)), which is \(2\times16\times4096\approx131{,}000\) numbers — about 0.8% as many — and still recover most of the quality of full fine-tuning. The scalar \(\alpha/r\) out front is just a knob for how strongly the add-on is applied.

Why such a small add-on is enough

The reason it works: adapting an already-trained model to a narrow task does not require rewriting what it knows, only nudging it in a particular direction. That nudge turns out to live in a small number of dimensions, so a thin, low-rank \(\Delta W\) is enough to capture it. You leave the base model's broad general ability completely untouched (frozen) and attach a small, swappable, task-specific add-on — often called an adapter. You can even keep several adapters for several tasks and load whichever one you need at the moment.

5.2 QLoRA — LoRA on top of a 4-bit base

QLoRA (Quantized LoRA) goes one step further. It loads the frozen base model in 4-bit NF4 (the 4-bit format from the last section), then trains the small LoRA adapters on top of that in full 16-bit precision. The big part — the base weights — is quartered in memory, while the small trainable part stays at high accuracy. This is what brings fine-tuning a 7B-to-14B model down to a single high-end consumer GPU such as an RTX 4090.

5.3 The four steps, in code

QLoRA fine-tuning on one GPU (Hugging Face stack)
# 1 · Load the base model in 4-bit NF4
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct",
                         quantization_config=bnb, device_map="auto")

# 2 · Attach the LoRA adapters (only these get trained)
from peft import LoraConfig, get_peft_model
model = get_peft_model(model, LoraConfig(
    r=16, lora_alpha=32, target_modules=["q_proj","v_proj"],
    task_type="CAUSAL_LM"))

# 3 · Train on your data (each example wrapped in the model's chat template!)
from trl import SFTTrainer
SFTTrainer(model=model, train_dataset=ds,
           dataset_text_field="text", max_seq_length=2048).train()

# 4 · Save just the tiny adapter (tens of MB, not tens of GB)
model.save_pretrained("./my_adapter")
Operator's implication — fine-tune for behavior, not for facts
A common and costly mistake is fine-tuning to stuff facts into the model ("teach it our whole product catalog"). Facts change, and they belong in a retrieval system (Part 5), where you can edit them without retraining anything. Fine-tune instead for behavior that is hard to get with a prompt: a fixed output format, a house writing style, a domain's way of reasoning, or better tool use. Then test the fine-tuned model against the golden set from Part 2 — fine-tuning without a way to measure the result is flying blind.

Part 6 · Open Models

Alignment: DPO & ORPO after SFT

Training a model by showing it good answers to copy is called supervised fine-tuning, or SFT. It teaches the model to imitate. There is a further step called preference tuning: instead of one right answer, you show the model two answers and tell it which one is better. This captures subtle wishes like "be helpful but keep it short" that are far easier to demonstrate by comparison than to spell out in a rule.

6.1 From RLHF to DPO

The classic way to do this, from the frontier post, is RLHF — Reinforcement Learning from Human Feedback. It first trains a separate model to score answers (a reward model), then uses reinforcement learning (the PPO method) to push the main model toward high-scoring answers. It works, but it is a fragile pipeline with several moving parts. DPO (Direct Preference Optimization) does the same job in one shot: given pairs of answers — a chosen one \(y_w\) and a rejected one \(y_l\) — it adjusts the model directly using a single, simple loss, with no separate reward model and no reinforcement-learning loop.

$$\mathcal{L}_{\text{DPO}}=-\log\sigma\!\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}-\beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)$$

In plain words: raise the model's probability of producing the chosen answer and lower its probability of producing the rejected one — but measure both against a frozen copy of the original model (written \(\pi_{\text{ref}}\), the reference) so the model does not drift too far from where it started. Here \(\pi_\theta\) is the model being trained, \(\beta\) controls how firmly it pulls, and \(\sigma\) is the logistic function that turns the score gap into a value between 0 and 1. ORPO (Odds-Ratio Preference Optimization) goes further still: it folds the preference step into ordinary SFT so it is one stage, with no reference copy needed at all.

The usual modern recipe
Base model → SFT (teach the task and the format using high-quality examples) → DPO or ORPO (refine using preferences about which answer is better). For tasks that need heavy reasoning, a reinforcement-learning stage with checkable rewards (GRPO, from Part 3) may come after. Each stage is smaller and cheaper than the one before, and each should have to pass the evaluation harness from Part 2 before you keep it.

6.2 The tooling

You almost never write these methods from scratch. Hugging Face PEFT is the library that provides LoRA and QLoRA adapters. Hugging Face TRL provides ready-made trainers for SFT, DPO, and GRPO. Unsloth is a library that makes training 2–5× faster while using less VRAM. Axolotl and LlamaFactory wrap the whole process so you drive it from a config file instead of code. The tooling is mature enough that the code is no longer the hard part — the hard parts are the data and the evaluation.

Appendix

References & further reading

Grouped by type; freely available online. Model specifics reflect public information as of early 2026 and change quickly — verify against official model cards.

Architecture & reasoning

  1. Jiang et al. (2024). Mixtral of Experts. Sparse MoE, 8 experts top-2. https://arxiv.org/abs/2401.04088
  2. Dai et al. (2024). DeepSeekMoE: Fine-grained & shared experts. https://arxiv.org/abs/2401.06066
  3. DeepSeek-AI (2024). DeepSeek-V3 Technical Report. 671B/37B MoE, aux-loss-free balancing, FP8. https://arxiv.org/abs/2412.19437
  4. DeepSeek-AI (2025). DeepSeek-R1. RL-trained reasoning. https://arxiv.org/abs/2501.12948
  5. Shao et al. (2024). DeepSeekMath (GRPO). https://arxiv.org/abs/2402.03300
  6. Fedus, Zoph & Shazeer (2021). Switch Transformers. Load-balancing loss. https://arxiv.org/abs/2101.03961

Quantization & fine-tuning

  1. Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. https://arxiv.org/abs/2106.09685
  2. Dettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NF4, double quant. https://arxiv.org/abs/2305.14314
  3. Frantar et al. (2022). GPTQ. https://arxiv.org/abs/2210.17323
  4. Lin et al. (2023). AWQ: Activation-aware Weight Quantization. https://arxiv.org/abs/2306.00978
  5. Rafailov et al. (2023). Direct Preference Optimization (DPO). https://arxiv.org/abs/2305.18290
  6. Hong et al. (2024). ORPO: Monolithic Preference Optimization. https://arxiv.org/abs/2403.07691

Tools

  1. Hugging Face. PEFT & TRL. https://huggingface.co/docs/peft
  2. Unsloth. Faster QLoRA fine-tuning. https://github.com/unslothai/unsloth

Model families, licenses, and specs are described at the level of well-established public understanding as of early 2026; proprietary training details are not disclosed and speculative items are flagged in-text.