Whetstone
Fine-tuning for small language models

Teach a small model one job.
Know it's better before you ship it.

Whetstone turns your own examples into a small, specialised language model. It checks the data before any GPU time is spent, puts people in charge at three clear decision points, and scores the result against the model you started with — so "fine-tuned" means improved, not just changed.

Try the live demoLive demo deployment isn't ready yet Read the documentation
The scorecard: the fine-tuned model compared metric by metric with its prompt-only baseline, with sample answers and the ship, iterate or abandon decision.
The scorecard at the third decision point: the fine-tune against its own baseline, with real sample answers, before anyone decides to ship. (Shown from a CPU practice run, which is labelled as such in the product.)
The problem

Fine-tuning is easy to start and hard to trust.

The training step itself has become a commodity. What goes wrong is everything around it.

Bad data trains silently

A garbled figure like "$12.SM" from a scanned document becomes something the model confidently repeats.

Duplicates, empty rows and made-up numbers rarely stop a training run.

Test results that lie

If the same company appears in practice and test examples, the model has seen the answers. Scores look great; real use doesn't.

Splitting by row instead of by entity is the usual culprit.

"Lower loss" isn't "better"

A falling training curve says the model fits its data — not that it beats simply prompting the original model.

Without a baseline, nobody can say whether the fine-tune was worth it.

Settings drift, runs don't reproduce

Library versions rename options; a run from last month can't be re-created because nobody recorded what it used.

Silent defaults are the enemy of reproducibility.

GPU compatibility surprises

A driver that's too old crashes training an hour in, with an error nobody can read.

The cheapest failure is the one caught before the GPU starts.

Nobody owns the decision

Scripts run end to end; a person only finds out what was trained after it's deployed.

Good tools make the human decisions explicit — and recorded.

The product

One path, three decisions, nothing hidden.

A fine-tune in Whetstone is one job that moves through a pipeline and pauses wherever a person should decide. Every step writes a record of exactly what it used.

Approve the data

Upload your examples and read a quality report. Suspicious numbers are flagged for a person to keep or send back to the source — never "fixed" automatically. Page through the built training examples and a split that keeps each company in one place.

Confirm the training setup

Pick a model and method in a guided wizard. Recommended settings are one click; every parameter the training libraries offer is there when you need it, checked before it can run. A GPU check says up front whether this machine can train.

Judge the result

A scorecard compares the fine-tune with its starting model on held-out examples, side by side, with invented numbers flagged. Ship it to the registry, iterate with the same settings pre-filled, or abandon it.

Architecture

Runs on one machine. Trains wherever the GPU is.

Everything ships as containers. The control plane needs no GPU; training, conversion and evaluation of real models are sent to a GPU host, checked for compatibility first.

Dashed: needs NVIDIA hardware. Without it, the whole path still runs as a clearly labelled CPU practice run.

Requirements vs. what's built

An honest feature matrix.

Every requirement we set, and where it stands today. "Partial" means the code is written and tested without a GPU, but hasn't yet run on real GPU hardware.

Built works and is tested end to end Partial written and tested, awaiting a GPU run Roadmap planned, not started
RequirementWhat was builtStatus
Data
Bring your own examplesUpload of JSON Lines files with size, encoding and per-line checks; provenance recordedBuilt
Know the data's quality firstQuality report: duplicates, diversity, length histogram, empty rows, outliersBuilt
Never trust a garbled numberSuspicious numbers flagged for a person to keep or send back; never alteredBuilt
Honest test dataSplit by entity with a leakage check; a dominant entity blocks approvalBuilt
Decisions
People approve the dataGate 1: built examples, split, re-split, approve or reject with a noteBuilt
People confirm the setupGate 2: guided wizard with GPU check, equivalent command line, explicit confirmBuilt
People judge the resultGate 3: scorecard verdict — ship, iterate (settings pre-filled) or abandonBuilt
Training
Train on a GPU with a choice of enginesHugging Face and Unsloth engines; Unsloth chosen by the model's architecture, with a fallbackPartial
Modern adapter methodsLoRA, QLoRA, DoRA, VeRA and LoRA+ — configurations validated against the real librariesPartial
Full control over settingsEvery trainer, adapter and loading parameter, generated from the pinned libraries, checked twiceBuilt
Try the pipeline without a GPUCPU practice run through every stage, labelled everywhere it appearsBuilt
Compute
Use our own machinesOn-premises Docker backend with live logs and clean teardownBuilt
Rent a GPU when neededRunPod: the key is saved and tested from Settings; running jobs on RunPod is next on the roadmapPartial
No driver surprisesPreflight check picks a compatible image or explains the fixBuilt
Reproducible GPU environmentsPinned worker images for both engines (written; not yet built on a GPU host)Partial
Evaluation & delivery
Prove it improvedBaseline-relative scorecard with metric directions, samples and invented-number flagsBuilt
Keep the test set fairTracks when a test set was first looked at and flags later runs judged on itBuilt
Versioned releasesRegistry packaging the adapter with every stage recordBuilt
Run it locally after trainingGGUF conversion and Ollama deploy (queued on a GPU worker; refused honestly without one)Partial
Try models immediatelyPlayground on the stack's own Ollama, with side-by-side comparisonBuilt
Platform
Scriptable as well as clickableAPI-first, with a command-line client at parityBuilt
Reproducible by constructionA record (manifest) from every stage and decision; pinned dependenciesBuilt
Secure by defaultKey sent only in headers, strict content policy, key proven absent from logsBuilt
Usable by everyoneAccessibility checks on every screen; the wizard works by keyboard aloneBuilt
Least privilege in containersRun the application containers as a non-root userRoadmap
Spend limitsBudget and runtime limits are recorded with each job but not yet enforcedRoadmap
Teams and sign-inMulti-user accounts and single sign-on (today: one deployment, one key)Roadmap
High-throughput servingvLLM behind the same serving interfaceRoadmap

28 requirements: 19 built, 5 partial (all awaiting GPU hardware), 4 on the roadmap.

Roadmap

What comes next.

Next

  1. A first run on real GPU hardware — build the pinned worker images and prove both engines, all five adapter methods, and GGUF conversion end to end. This turns every "Partial" above into "Built".
  2. Non-root containers — files written by the stack owned by the operator, not by root.
  3. Enforced limits — stop a run at its budget or runtime cap.

Later

  1. Teams — accounts, sign-in and per-project access.
  2. vLLM serving for higher throughput.
  3. More compute backends behind the same interface.
  4. A cloud "twin" demo — the same application on a small cloud box, no install needed.
Deployment

Your data and your models stay on your hardware.

The whole stack — database, queue, API, worker, web front and model server — starts with one command on any Docker host. Real training goes to a GPU host you control.

# start the stack, wait until healthy, migrate, pull the default serving model
scripts/stack_up.sh

# optional: a walkable demo (a shipped run, and a job waiting for its first decision)
WHET_API_URL=http://localhost:8080/api WHET_API_KEY=<key> uv run python scripts/demo_seed.py

Upgrades

Database changes ship as migrations; alembic upgrade head brings an existing install up to date, and every migration can be rolled back.

Configuration

Everything is an environment variable — the access key, upload limits, the serving model, and where artifacts live. Secrets never pass through the browser.

GPU hosts

Pinned worker images per engine and CUDA version; a preflight check matches the host's driver before any work starts.

Live demo

Read the story. Then try the real thing.

The demo is the actual product, seeded with made-up companies. One fine-tune has already been shipped; another is waiting for its first decision — yours. Access is by invitation: ask us for a demo key. (The public demo is being set up.)

Open the live demoLive demo deployment isn't ready yet How to use it