← All writing
Open modelsLab deep dive

The model weights are open. Can we see how the model learned?

Dhruvi Paprunia8 min read
Open weights beside an open training notebook

Why open model training matters, and how to test the steps that shape an assistant.

You can download a model and still have no clear idea which training step taught it a behavior. Imagine an assistant that finally gives the right answer but stops following your request to keep it short.

Explained simply

An AI can get better at answering questions and worse at following instructions.

Seeing the examples it learned from, the training code and earlier versions helps researchers investigate what changed.

Sharing that history makes the AI easier to study, not automatically better.

Open weights are not the whole explanation.

An AI model is a program trained on examples so it can respond to a question. Its weights are the numbers that store what it learned. Releasing those weights lets someone run the model.

To study why its behavior changed, we also need the training data, code and saved versions along the way.

That is why Ai2's work interests me. The Allen Institute for AI is an independent nonprofit research institute.

They connect, but they are not three names for one model.

What you'll learn

These notes draw on Tülu 3 by Nathan Lambert, Jacob Morrison, Valentina Pyatkin and their co-authors, and 2 OLMo 2 Furious by the OLMo Team, with Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo and their co-authors.

The useful contribution is an inspectable training recipe, not a promise that every open model behaves better.

How does a base model become an assistant?

A base model has learned patterns by predicting the next piece of text, but it may not follow a request well. The next phase, called post-training, teaches it to respond in ways people can use.

Tülu 3 describes three stages. Each stage gives the model a different kind of feedback.

A prompt is the question or instruction given to a model. First is supervised fine-tuning, or SFT. The model studies a prompt beside an example answer.

Suppose the prompt says, "Explain this result in two sentences." The example shows a clear answer in two sentences. The model learns from many such pairs rather than discovering the desired response by itself.

The mix of examples matters. If most examples are math problems, the model might improve at math without improving at a very different task. Ai2 describes how it selects, cleans, and mixes training data, so a researcher can ask what changed instead of looking only at the final score.

Next is Direct Preference Optimization, or DPO. Now the feedback compares two answers to the same prompt and marks one as preferred.

Imagine two answers that both solve a problem, but only one includes the explanation the user asked for. The comparison says more than a single example answer: it shows the model which difference mattered.

Some of the answers used for Tülu's comparisons came from the model being trained. This is called on-policy data. It matters because that model's own mistakes may differ from another model's mistakes.

If all comparisons come from someone else's answers, the training may miss exactly the errors this model makes.

The final stage is reinforcement learning with verifiable rewards, or RLVR. The model tries an answer and receives a reward when a checker can confirm the result.

For a math problem with a known answer, the checker can test the final number. If the task is "What is 7 + 5?", a simple illustration is reward = 1 when the final answer is 12, and 0 otherwise. These are example reward values, not the full reward settings reported in the paper.

A precise instruction can also have a checkable condition. The model then updates using that feedback.

A correct final number does not prove that its explanation makes sense. A thoughtful research idea or a helpful refusal is harder to mark right or wrong with a simple check. RLVR works best when the checker really matches the task. That limit matters as much as the method itself.

Illustrative feedback, not measured results or the full training reward.

Illustrative feedback, not measured results or the full training reward.

Tülu reward learning method diagram

Figure 18 from Tülu 3 by Nathan Lambert and co-authors. The original paper starts with Llama 3.1. Prompts are instructions, completions are generated answers, policy means the model, and scalar means one reward number. The Greek letter gamma is the reward for a correct answer; the update equation changes the model weights using that feedback. This is a method diagram, not an accuracy result.

Where do OLMo and Tülu meet?

One distinction is easy to lose: the models in the original Tülu 3 paper start with Llama 3.1, a model family released by Meta, as their base. Tülu is the training recipe and the family of models produced with it. It is not the name of OLMo 2's earlier training.

That earlier training is pretraining: Ai2 started from scratch and taught OLMo 2 to predict the next piece of text using a large data mix that it released. A shorter final phase used a smaller, higher-quality mix to strengthen skills such as math.

This produced the base model. The Tülu recipe then trained it to act as an assistant.

Ai2 later applied the Tülu 3 recipe to OLMo 2. The OLMo release includes information about its data, code, training method, weights, and checkpoints. A checkpoint is a saved version of the model at a particular point in training.

Ai2 also adapted parts of the later training, including preference comparisons generated from OLMo's own answers. Applying a recipe to a new base was not just a copy and paste.

The paper figure below shows scores during OLMo 2's later reinforcement learning runs. A score summarizes performance on a set of tests, called evaluations.

The changing lines are useful precisely because they show that the outcome depends on the training path, not only the final model.

Figure 13 detail from 2 OLMo 2 Furious by the OLMo Team. Average and IFEval scores during reward-based training. Separate colors are separate runs, not the three post-training stages.

Figure 13 detail from 2 OLMo 2 Furious by the OLMo Team. Average and IFEval scores during reward-based training. Separate colors are separate runs, not the three post-training stages.

With access to checkpoints, a researcher can compare a base model with versions saved after the example, preference, and reward stages. They can change one data mix or training goal and see what happens.

They can also check whether a test question appeared in training, which would make a good score less convincing. A finished set of weights alone cannot answer these questions.

A test card: did the model improve, or did the score hide a tradeoff?

This is a proposed experiment, not a result reported by Ai2. Start from the same model and run three training conditions:

Change only one setting in each experimental run.

Use the same held-out questions, meaning questions reserved for testing rather than training. Use the same checkable instruction tasks for all three, with no test examples added to training.

For each question, save whether the answer is correct and whether it follows the instruction.

Keep fixed: starting model, held-out questions, evaluation rules. Change one thing: harder problems OR more attempts. Record separately: wrong to right, right to wrong, instruction failures, computing cost.

Proposed test card. No gains or results are claimed.

What I would test next

The test card is a starting point for a question the open recipe makes possible: does harder reward training help the model discover new solutions, or mostly improve questions it nearly solved already? I would inspect which individual questions changed, not just the average.

More answer attempts also cost more computing power, so I would compare gains at a matched budget rather than call a more expensive run a better training method.

I would then test whether those math gains survive questions written differently and instructions that a simple checker cannot fully judge. A response can be correct and still be confusing or unhelpful.

That would need additional review, not just the same math reward. These are tests I would propose, not conclusions the papers establish.

Openness has limits.

The papers report average scores on sets of test questions. Our proposed test would:

The aim is to measure what improved, what regressed, and at what cost, not just whether the average rose.

The useful takeaway: an open final model lets us run it. An open training path lets us ask what changed it.

Sources: Tülu 3 paper · Tülu 3 technical write-up · OLMo 2 paper · OLMo 2 release · Open Instruct code · Ai2