This guide presents a fully public and low-cost method for significantly improving small models’ ability to follow structured output requirements. We fine-tune LFM2.5-350M using Group Relative Policy Optimization (GRPO) from the TRL library, and evaluate it on the IFStruct benchmark. The complete run requires only about 500 samples and 100 training steps—small enough to run on the free tier of Colab or on a Kaggle GPU. The full workflow is published on GitHub. The results show that even with a lightweight fine-tuning pipeline, the model’s performance on the IFStruct benchmark can improve from 22.6% to 29.7%.
Structured output is one of the most common real-world tasks for large language models (LLMs), yet most benchmarks include it as part of broader reasoning or information extraction scores rather than measuring it separately. Whether a model can reliably return valid, parseable output in the required format and structure—in other words, its schema compliance—often determines whether it can actually be integrated into downstream systems.
Note that the training workflow described here is not the workflow used to train the reinforcement learning models discussed in the IFStruct blog post. This notebook is not intended to reproduce the IFStruct benchmark score. Instead, it demonstrates how fine-tuning a smaller model for a specific task can improve its performance to a level comparable to that of much larger models.
Prerequisites
This guide is divided into two parts, each running in a different environment:
-
Fine-tuning requires a GPU. The accompanying notebook is sized to run on the free tier of Colab or on a Kaggle GPU.
-
Evaluation can be run locally on a MacBook using
llama.cpp(this article uses a MacBook Pro with an Apple M5 Max chip and 36 GB of unified memory).llama.cppprovides an OpenAI-compatible server, which the IFStruct evaluator communicates with.
We use uv for Python tool management and llama.cpp to serve the model. Following Liquid AI’s llama.cpp deployment documentation, install llama.cpp via Homebrew and verify that llama-server is available:
brew install llama.cppllama-server --versionEvaluating IFStruct on LFM2.5-350M (Base Model)
First, let’s evaluate LFM2.5-350M on the IFStruct benchmark and see whether we can reproduce the reported score of 21.1%.
IFStruct is a benchmark for testing the validity and schema-following capabilities of large language model outputs. The benchmark is open-sourced in Liquid4All/ifstruct, while the publicly available benchmark dataset is hosted on Hugging Face at LiquidAI/ifstruct-v1.0.
git clone https://github.com/Liquid4All/ifstruct.gitFor a fair evaluation comparison, we serve the model locally on a MacBook using llama.cpp. Here, we use the BF16 GGUF model (LiquidAI/LFM2.5-350M-GGUF).
Then run the following command to start the base model server:
llama-server \ -hf LiquidAI/LFM2.5-350M-GGUF:BF16 \ -c 32768 \ -np 4 \ -ngl 99 \ --alias LiquidAI/LFM2.5-350M \ --host 127.0.0.1 \ --port 8080-
--alias: The model name used by IFStruct when sending requests to the OpenAI-compatible endpoint -
-ngl 99: Instructsllama.cppto offload all layers to the GPU whenever possible -
-np 4: Processes 4 requests in parallel -
-c 32768: The prompt context size
Once the server is running, we can run the full benchmark with 2,000 samples:
uv run ifstruct-eval \ --model LiquidAI/LFM2.5-350M \ --base-url http://localhost:8080/v1 \ --api-key dummy \ --dataset data/test.jsonl \ --results-file results/lfm2.5-350m-llamacpp-base.json \ --n-threads 4 \ --max-tokens 2048 \ -v============================================================Model: LiquidAI/LFM2.5-350M============================================================Overall: 452/2000 passed (22.6%)Average latency: 1453ms
By format: JSON: 180/1000 passed (18.0%) YAML: 272/1000 passed (27.2%)
By top-level structure: Wrapper key 288/1011 passed (28.5%) Bare list 164/989 passed (16.6%)
By entity type: test__camera_review 6/83 passed (7.2%) test__clinical_trial 20/104 passed (19.2%) test__conference_schedule 7/87 passed (8.0%) test__escaping__bug_report_batch 24/89 passed (27.0%) test__escaping__config_snippet_audit 15/85 passed (17.6%) test__escaping__customer_email_thread 5/73 passed (6.8%) test__escaping__dialogue_sample 14/95 passed (14.7%) test__escaping__interview_transcript_segment 21/80 passed (26.2%) test__escaping__log_parser_examples 21/72 passed (29.2%) test__escaping__pr_discussion 22/87 passed (25.3%) test__escaping__repro_steps_batch 16/73 passed (21.9%) test__escaping__screenplay_scene 16/92 passed (17.4%) test__escaping__short_story_chapter 15/84 passed (17.9%) test__escaping__support_ticket_batch 27/73 passed (37.0%) test__escaping__terminal_session_notes 20/70 passed (28.6%) test__event_ticket_booking 49/107 passed (45.8%) test__gpu_review 6/94 passed (6.4%) test__invoice 28/86 passed (32.6%) test__job_posting 25/85 passed (29.4%) test__real_estate_listing 31/82 passed (37.8%) test__recipe 3/70 passed (4.3%) test__rental_car_booking 27/79 passed (34.2%) test__scientific_experiment 13/69 passed (18.8%) test__travel_itinerary 21/81 passed (25.9%)
Common errors: 7228x required field missing 738x wrong item count 540x type mismatch 317x Unclosed code block 190x extraneous field 'notes' 181x extraneous field 'path' 175x extraneous field 'constraints' 170x extraneous field 'type' 170x missing code block 100x expected bare list, got wrapperThe IFStruct release blog post reports a score of 21.1% for LFM2.5-350M. Using llama.cpp locally with the BF16 configuration, we measured 22.6%, which is close to the 21.1% reported in the IFStruct blog post. We use this local result as the baseline for comparisons within the same serving stack.
GRPO Fine-Tuning for Structured Output with TRL
The complete, runnable workflow is available in the accompanying notebook. This section covers only the relevant parts.
Training Data
We use the nvidia/Nemotron-RL-instruction_following-structured_outputs dataset, where each prompt corresponds to a target JSON Schema and an expected number of fields. Training uses approximately 500 samples.
Because the Nemotron data distribution differs from that of the IFStruct evaluation data, we augment the prompts to bridge two gaps between them:
-
40% of the samples append an instruction to “return the output in a fenced code block,” teaching the model to follow formatting instructions rather than always outputting raw JSON.
-
A separate, non-overlapping 20% of the samples are converted into top-level array tasks (wrapping the Schema in an
arrayand requiring a specified number of elements) to train bare-list output and item-count compliance.
Model and LoRA
We load LiquidAI/LFM2.5-350M and attach a LoRA adapter. Because LFM2.5 uses a hybrid attention/convolution architecture, we target LFM-specific module names:
lora_config = LoraConfig( r=16, lora_alpha=32, bias="none", task_type="CAUSAL_LM", target_modules=[ "q_proj", "k_proj", "v_proj", "out_proj