source: Hugging Face Blog: Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
level: technical
researchers fine-tuned liquid ai's lfm2.5-350m model using group relative policy optimization (grpo) with the trl library. the goal was to improve structured output compliance, measured by the ifstruct benchmark. the training used about 500 samples and 100 steps, small enough for a free-tier colab or kaggle gpu. the fine-tuned model was evaluated locally on a macbook via llama.cpp. the recipe is fully public and inexpensive, aiming to make small models more reliable at returning valid, parseable output in requested formats.
the base model scored 22.6% on ifstruct, close to the reported 21.1%. after grpo fine-tuning, the score rose to 29.7%, a gain of 7.1 percentage points. json pass rate jumped from 18.0% to 31.9%, while yaml stayed nearly flat at 27.2% to 27.5%. bare list outputs improved from 16.6% to 29.7%. common errors remained, such as missing required fields and wrong item counts, but the fine-tuned model reduced some format-specific mistakes. the result is still below qwen3.5-2b's 33.15%, but the gap narrowed.
the training used a lora adapter targeting about 6 million parameters, roughly 1.66% of the model. three reward functions scored json format, field count, and schema validation, with weights 1.0, 0.5, and 2.0. the prompt data came from nvidia's nemotron structured outputs dataset, augmented to include fenced code blocks and top-level arrays. this approach shows that a cheap, task-specific reward signal can make a small model substantially more reliable about output form, closing much of the gap to larger models.
why it matters: reliable structured output is essential for wiring language models into downstream systems, and this recipe shows a low-cost way to improve it.
source: Hugging Face Blog: Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps