Résumé

LLM fine-tuning

A 3-billion-parameter language model, fine-tuned to solve grade-school math word problems in the voice of a cheerful plumber. It runs live below.

Result
The persona didn’t cost accuracy. Fine-tuning did.
My role
I owned the training: supervised fine-tuning (SFT), then an RLAIF follow-up in which an AI judge ranked the model’s own answers to build 773 preference pairs for DPO. That raised the judge’s persona score from 2.58 to 2.93 out of 5, with no clear change in accuracy.
Context
Harvard AI Safety (CS 2881) (opens in a new tab) Team of three, 2026
Built with
Python, PyTorch, TRL, PEFT (LoRA), llama.cpp

Try it

Or start from
Mario’s working

Nothing yet. Ask a question and the steps appear here as the model writes them.

It gets about one problem in three wrong. Check its work.

What we found

Fine-tuning made the model sound more like Mario than prompting did. It also made it worse at math, and so did fine-tuning without Mario.

500 held-out GSM8K test questions. Voice similarity is character n-gram cosine similarity to held-out Mario reference answers, which is a rough proxy and not a human rating.
ModelTest accuracy95% intervalVoice similarity
Original Qwen2.5-3B73.0%69.0–76.80.356
Original, told to be Mario77.6%74.0–81.20.359
Fine-tuned on plain solutions65.2%61.0–69.40.346
Fine-tuned on Mario solutions66.6%62.4–70.60.402

The highlighted row is the model running on this page. The two fine-tuned models lost about the same accuracy, so the drop can’t be pinned on the persona. It looks like a cost of fine-tuning on 3,000 examples at all.

How it was made

  1. Step 1, data 3,000 GSM8K solutions rewritten in Mario’s voice Every rewrite was checked automatically to confirm it still ended on the correct final answer.
  2. Step 2, training LoRA fine-tune of Qwen2.5-3B-Instruct Rank 16 on all attention and MLP projections, with no system prompt, so the persona lives in the weights.
  3. Step 3, evaluation Four models, same 500 test questions A plain-solution fine-tune served as the control, to separate the cost of the persona from the cost of fine-tuning.
  4. Step 4, serving 4-bit weights on a free shared CPU Adapter merged, quantized to 1.9 GB and served with llama.cpp. One question at a time, and it sleeps when nobody visits.

Made with two teammates for Harvard’s AI Safety course (CS 2881). Built with Qwen under the Qwen Research License, for non-commercial demonstration. Mario is a Nintendo character; this project is not affiliated with Nintendo.

Talk to me

I’m glad to walk through the evaluation design, or what I’d test next.

Other projects