← All writing
Reinforcement Learning

The polite sentence that broke my model

How we taught small language models to call tools with GRPO, and why the strictest rule worked best

Dhruvi Paprunia6 min read · Also on Medium ↗
Explainer: GRPO took Qwen2.5-3B from 6.1% to 71.1% tool-call accuracy on a single T4 GPU
The whole paper on one page.

"This is the correct tool call:"

That was the problem. Six polite words, followed by a perfectly good JSON object.

To a human, that answer looks fine. To a program waiting for JSON, it is garbage. The parser breaks, the tool never runs, and the agent fails.

This is the story of how my co-author Vansh Kharidia and I, with our advisor Dr. Pankti Doshi, trained small language models to stop doing that, and to pick the right tool with the right arguments. On a single T4 GPU, Qwen2.5-3B went from 6.1% to 71.1% accuracy on tool calls.

First, what is tool use?

A language model on its own can only produce text. It can't check today's weather, run a calculation it isn't sure about, or look something up in a database.

Tool use fixes that. You give the model a list of tools (functions, APIs), and when it needs one, it writes a structured call, usually in JSON:

[{"name": "qr_code_image", "arguments": {"size": 7, "url": "example.com"}}]

Your code reads that JSON, runs the real function, and hands the result back. This is how AI agents get things done in the real world.

Big models are pretty good at this. Small language models (SLMs), roughly 100M to 5B parameters, are not. And small models are exactly the ones you want on a phone, on an edge device, or anywhere compute is tight.

The gap we saw

We tested three small models on tool calls from Salesforce's xLAM function-calling dataset: Qwen2.5-1.5B, Qwen2.5-3B and Llama 3.2 3B. We trained on 4,000 examples and tested on 1,000.

Before any training, Qwen2.5-3B wrote valid JSON 100% of the time. Its overall accuracy was 6.1%.

So why only 6.1%, if the JSON was always valid?

Because valid JSON is the easy part. The paper counts an answer as right only when three things are right at the same time: the JSON, the tool name, and the arguments. One slip anywhere and the whole call counts as wrong.

The paper doesn't break the 6.1% down mistake by mistake. But while training, we saw the same slips again and again:

Here is a real example from the paper. The task needed a QR code tool called qr_code_image with a size and a URL.

A small mistake like that means nothing to a person. For a program, it means total failure.

Why reinforcement learning, and what GRPO is

The usual fix is supervised fine-tuning: show the model thousands of correct answers and have it copy them. We wanted something that teaches the model what a good answer is, not just what one looks like.

That is reinforcement learning (RL). The model tries, a reward function scores the attempt, and the model shifts toward whatever scored higher.

The classic RL method for language models is PPO. PPO needs a second network, called a critic, to estimate how good each answer should be. The critic is usually about as big as the model you are training, and on a single T4 that is a problem.

GRPO (Group Relative Policy Optimization) drops the critic. It was introduced in the DeepSeekMath paper and became famous through DeepSeek-R1. The idea is simple:

  1. Give the model one question.
  2. Let it write several answers. We used 8.
  3. Score every answer with the reward function.
  4. Compare each answer to the group's average. Above average gets pushed up, below average gets pushed down.

The group is its own baseline. No critic, so a lot less memory. That is what made this work on one GPU.

Diagram of GRPO: one question, eight scored answers, compared with the group average
How GRPO learns. Example scores.

GRPO is only as good as its reward, though. So most of our work went into the reward.

Designing the reward

We wanted to reward three things: valid JSON, the right tool name, and the right arguments. The final version gives up to 1.0 point per answer:

And two rules sit on top of that:

Reward: zero for any extra text, otherwise 0.125 for JSON, 0.375 for the tool name and 0.5 for the arguments
Our final reward, from the paper.

We didn't start with these numbers. We got to them by watching the model learn.

A reward that changes as the model learns

Early on, valid JSON was worth 0.5. The models picked up the format fast, so JSON stopped being the problem. We cut its weight to 0.125 so learning could shift to the harder part: choosing tools and filling in arguments.

Then we trained on tool names and arguments separately to see which was harder. Both were hard, but the mistakes were different. Tool names stayed fairly stable. Argument names were where the models kept making small hallucinations, like a slightly wrong key. So arguments got the biggest share, 0.5.

We call this capability-aware reward modeling. You don't design the perfect reward on day one. You watch what the model already knows, and move the reward to what it doesn't.

What didn't work: being gentle

Before the strict version, we tried a softer reward. Small mistakes got small penalties instead of zero. It sounds fairer. It worked worse, in three ways:

Soft penalties let the model find a comfortable middle. Strict ones leave it only one way to score.

Our "aha" moment

Once any text outside the JSON meant zero reward, the models dropped the chatter fast.

You can see it in the training logs. The average answer got much shorter in the first stretch of training, most sharply for Llama 3.2 3B. The model learned that the only safe answer was the JSON and nothing else.

Training chart: completion length falls early in training
Completion length during training (Fig. 2 in the paper).

The reward rose steadily too. Here is Qwen2.5-1.5B over 500 training steps:

Training chart: reward rises for Qwen2.5-1.5B over 500 steps
Reward during training, Qwen2.5-1.5B (Fig. 3 in the paper).

The results

After 500 training steps with GRPO and LoRA (rank 64) on a single T4 GPU:

Bar chart: tool-call accuracy before and after GRPO for three models
Results from the paper.

Every model ended at 100% valid JSON. Accuracy went up for all three, and the 3B Qwen got the most out of training.

These numbers come with limits. We trained for only 500 steps because of compute. We tested on 1,000 examples from one dataset. We didn't compare GRPO head-to-head against other RL methods. Those are the next questions, along with messier real-world APIs.

What I took away

I observed three things that I now bring to every project.

Strict beats clever. Our fanciest rewards gave the model room to cheat. A simple zero for extra text did more than any clever partial credit.

Watch the model, then change the reward. The best reward wasn't the one we wrote first. It came from seeing what the model had already mastered and moving the pressure elsewhere.

Small models can do more than we think. Nothing about Qwen2.5-3B changed except how we rewarded it, and its accuracy went from 6.1% to 71.1%.

My takeaway: small models often don't need to get bigger. They need a clearer idea of what counts as a right answer.


Read the full paper: Advancing SLM Tool-Use Capability using Reinforcement Learning (IEEE AIC), by Dhruvi Paprunia, Vansh Kharidia and Dr. Pankti Doshi.

Sources: Our paper on arXiv · GRPO, from DeepSeekMath (Shao et al.) · DeepSeek-R1 · PPO (Schulman et al.) · xLAM function-calling dataset

Dhruvi Paprunia6 min read · Also on Medium ↗
More writing →