Teach a small model one job.
Know it's better before you ship it.
Whetstone turns your own examples into a small, specialised language model. It checks the data before any GPU time is spent, puts people in charge at three clear decision points, and scores the result against the model you started with — so "fine-tuned" means improved, not just changed.
Fine-tuning is easy to start and hard to trust.
The training step itself has become a commodity. What goes wrong is everything around it.
Bad data trains silently
A garbled figure like "$12.SM" from a scanned document becomes something the model confidently repeats.
Duplicates, empty rows and made-up numbers rarely stop a training run.
Test results that lie
If the same company appears in practice and test examples, the model has seen the answers. Scores look great; real use doesn't.
Splitting by row instead of by entity is the usual culprit.
"Lower loss" isn't "better"
A falling training curve says the model fits its data — not that it beats simply prompting the original model.
Without a baseline, nobody can say whether the fine-tune was worth it.
Settings drift, runs don't reproduce
Library versions rename options; a run from last month can't be re-created because nobody recorded what it used.
Silent defaults are the enemy of reproducibility.
GPU compatibility surprises
A driver that's too old crashes training an hour in, with an error nobody can read.
The cheapest failure is the one caught before the GPU starts.
Nobody owns the decision
Scripts run end to end; a person only finds out what was trained after it's deployed.
Good tools make the human decisions explicit — and recorded.
One path, three decisions, nothing hidden.
A fine-tune in Whetstone is one job that moves through a pipeline and pauses wherever a person should decide. Every step writes a record of exactly what it used.
Approve the data
Upload your examples and read a quality report. Suspicious numbers are flagged for a person to keep or send back to the source — never "fixed" automatically. Page through the built training examples and a split that keeps each company in one place.
Confirm the training setup
Pick a model and method in a guided wizard. Recommended settings are one click; every parameter the training libraries offer is there when you need it, checked before it can run. A GPU check says up front whether this machine can train.
Judge the result
A scorecard compares the fine-tune with its starting model on held-out examples, side by side, with invented numbers flagged. Ship it to the registry, iterate with the same settings pre-filled, or abandon it.
Data you can see into
Quality tiles, a length histogram, lint checks, and a review queue for every suspicious number.
A split that can't leak
Groups stay in one split; a group that dominates the data blocks approval until it's addressed.
Every setting, none by surprise
Your changes stay on top with undo; everything else shows its default. Only what you change is saved.
Watch it happen
Stages and a live log as the work runs; each waiting decision links straight to its screen.
Versions, not folders
Shipped models are packaged with every record of how they were made.
Try it right away
Ask any served model, optionally side by side, with invented numbers flagged the same way as the scorecard.
Runs on one machine. Trains wherever the GPU is.
Everything ships as containers. The control plane needs no GPU; training, conversion and evaluation of real models are sent to a GPU host, checked for compatibility first.
Dashed: needs NVIDIA hardware. Without it, the whole path still runs as a clearly labelled CPU practice run.
An honest feature matrix.
Every requirement we set, and where it stands today. "Partial" means the code is written and tested without a GPU, but hasn't yet run on real GPU hardware.
| Requirement | What was built | Status |
|---|---|---|
| Data | ||
| Bring your own examples | Upload of JSON Lines files with size, encoding and per-line checks; provenance recorded | Built |
| Know the data's quality first | Quality report: duplicates, diversity, length histogram, empty rows, outliers | Built |
| Never trust a garbled number | Suspicious numbers flagged for a person to keep or send back; never altered | Built |
| Honest test data | Split by entity with a leakage check; a dominant entity blocks approval | Built |
| Decisions | ||
| People approve the data | Gate 1: built examples, split, re-split, approve or reject with a note | Built |
| People confirm the setup | Gate 2: guided wizard with GPU check, equivalent command line, explicit confirm | Built |
| People judge the result | Gate 3: scorecard verdict — ship, iterate (settings pre-filled) or abandon | Built |
| Training | ||
| Train on a GPU with a choice of engines | Hugging Face and Unsloth engines; Unsloth chosen by the model's architecture, with a fallback | Partial |
| Modern adapter methods | LoRA, QLoRA, DoRA, VeRA and LoRA+ — configurations validated against the real libraries | Partial |
| Full control over settings | Every trainer, adapter and loading parameter, generated from the pinned libraries, checked twice | Built |
| Try the pipeline without a GPU | CPU practice run through every stage, labelled everywhere it appears | Built |
| Compute | ||
| Use our own machines | On-premises Docker backend with live logs and clean teardown | Built |
| Rent a GPU when needed | RunPod: the key is saved and tested from Settings; running jobs on RunPod is next on the roadmap | Partial |
| No driver surprises | Preflight check picks a compatible image or explains the fix | Built |
| Reproducible GPU environments | Pinned worker images for both engines (written; not yet built on a GPU host) | Partial |
| Evaluation & delivery | ||
| Prove it improved | Baseline-relative scorecard with metric directions, samples and invented-number flags | Built |
| Keep the test set fair | Tracks when a test set was first looked at and flags later runs judged on it | Built |
| Versioned releases | Registry packaging the adapter with every stage record | Built |
| Run it locally after training | GGUF conversion and Ollama deploy (queued on a GPU worker; refused honestly without one) | Partial |
| Try models immediately | Playground on the stack's own Ollama, with side-by-side comparison | Built |
| Platform | ||
| Scriptable as well as clickable | API-first, with a command-line client at parity | Built |
| Reproducible by construction | A record (manifest) from every stage and decision; pinned dependencies | Built |
| Secure by default | Key sent only in headers, strict content policy, key proven absent from logs | Built |
| Usable by everyone | Accessibility checks on every screen; the wizard works by keyboard alone | Built |
| Least privilege in containers | Run the application containers as a non-root user | Roadmap |
| Spend limits | Budget and runtime limits are recorded with each job but not yet enforced | Roadmap |
| Teams and sign-in | Multi-user accounts and single sign-on (today: one deployment, one key) | Roadmap |
| High-throughput serving | vLLM behind the same serving interface | Roadmap |
28 requirements: 19 built, 5 partial (all awaiting GPU hardware), 4 on the roadmap.
What comes next.
Next
- A first run on real GPU hardware — build the pinned worker images and prove both engines, all five adapter methods, and GGUF conversion end to end. This turns every "Partial" above into "Built".
- Non-root containers — files written by the stack owned by the operator, not by root.
- Enforced limits — stop a run at its budget or runtime cap.
Later
- Teams — accounts, sign-in and per-project access.
- vLLM serving for higher throughput.
- More compute backends behind the same interface.
- A cloud "twin" demo — the same application on a small cloud box, no install needed.
Your data and your models stay on your hardware.
The whole stack — database, queue, API, worker, web front and model server — starts with one command on any Docker host. Real training goes to a GPU host you control.
# start the stack, wait until healthy, migrate, pull the default serving model
scripts/stack_up.sh
# optional: a walkable demo (a shipped run, and a job waiting for its first decision)
WHET_API_URL=http://localhost:8080/api WHET_API_KEY=<key> uv run python scripts/demo_seed.py
Upgrades
Database changes ship as migrations; alembic upgrade head brings an existing install up to date, and every migration can be rolled back.
Configuration
Everything is an environment variable — the access key, upload limits, the serving model, and where artifacts live. Secrets never pass through the browser.
GPU hosts
Pinned worker images per engine and CUDA version; a preflight check matches the host's driver before any work starts.
Read the story. Then try the real thing.
The demo is the actual product, seeded with made-up companies. One fine-tune has already been shipped; another is waiting for its first decision — yours. Access is by invitation: ask us for a demo key. (The public demo is being set up.)