LLM fine-tuning
A 3-billion-parameter language model, fine-tuned to solve grade-school math word problems in the voice of a cheerful plumber. It runs live below.
- Result
- The persona didn’t cost accuracy. Fine-tuning did.
- My role
- I owned the training: supervised fine-tuning (SFT), then an RLAIF follow-up in which an AI judge ranked the model’s own answers to build 773 preference pairs for DPO. That raised the judge’s persona score from 2.58 to 2.93 out of 5, with no clear change in accuracy.
- Context
- Harvard AI Safety (CS 2881) (opens in a new tab) Team of three, 2026
- Built with
- Python, PyTorch, TRL, PEFT (LoRA), llama.cpp
Try it
Nothing yet. Ask a question and the steps appear here as the model writes them.
It gets about one problem in three wrong. Check its work.
What we found
Fine-tuning made the model sound more like Mario than prompting did. It also made it worse at math, and so did fine-tuning without Mario.
| Model | Test accuracy | 95% interval | Voice similarity |
|---|---|---|---|
| Original Qwen2.5-3B | 73.0% | 69.0–76.8 | 0.356 |
| Original, told to be Mario | 77.6% | 74.0–81.2 | 0.359 |
| Fine-tuned on plain solutions | 65.2% | 61.0–69.4 | 0.346 |
| Fine-tuned on Mario solutions | 66.6% | 62.4–70.6 | 0.402 |
The highlighted row is the model running on this page. The two fine-tuned models lost about the same accuracy, so the drop can’t be pinned on the persona. It looks like a cost of fine-tuning on 3,000 examples at all.
How it was made
- Step 1, data 3,000 GSM8K solutions rewritten in Mario’s voice Every rewrite was checked automatically to confirm it still ended on the correct final answer.
- Step 2, training LoRA fine-tune of Qwen2.5-3B-Instruct Rank 16 on all attention and MLP projections, with no system prompt, so the persona lives in the weights.
- Step 3, evaluation Four models, same 500 test questions A plain-solution fine-tune served as the control, to separate the cost of the persona from the cost of fine-tuning.
- Step 4, serving 4-bit weights on a free shared CPU Adapter merged, quantized to 1.9 GB and served with llama.cpp. One question at a time, and it sleeps when nobody visits.
Made with two teammates for Harvard’s AI Safety course (CS 2881). Built with Qwen under the Qwen Research License, for non-commercial demonstration. Mario is a Nintendo character; this project is not affiliated with Nintendo.
Talk to me
I’m glad to walk through the evaluation design, or what I’d test next.