Whetstone ← Overview

Documentation

How to use Whetstone, what happens underneath, and how to drive it from code.

1. User manual

A fine-tune is a job. It moves through a pipeline on its own and pauses at three gates, where a person decides. The dashboard's Needs your attention list shows every job waiting for you, with a Review button that opens the right screen.

Signing in (demo)

  1. Open the demo address you were given and enter the demo key (sent with your invitation).
  2. The key stays in this browser tab only and is sent with each request as a header — never in a web address.
  3. The demo is seeded with made-up companies in the project acme-pitches (demo): one fine-tune has already been shipped, and one is waiting at gate 1 for you.
The demo runs without a GPU. Training is a CPU practice run: every stage runs and every screen fills in, but the model isn't really trained, and its scores are placeholders. The product labels this wherever it appears.

Datasets & quality

Choose a project and drop a .jsonl file (one example per line, with fields such as input/output and optionally group_id for the company or document it belongs to). You'll see quality tiles, a length histogram, lint checks and the number review queue.

Datasets screen with the quality report and the number-review queue.

For each flagged number, choose Keep (it's correct) or Edit source (it's wrong; add a note). Whetstone never changes your file: a number marked for editing blocks gate 1 until a corrected file is uploaded. When you're ready, press Build & split.

Gate 1 · Build & split

Page through the examples exactly as the model will see them, and check the split into training, validation and test sets. Each company stays in one set, so the test can't be learned by heart.

Build and split review screen.

Gate 2 · New fine-tune

Four steps: Dataset (jobs that passed gate 1), Model & method, Compute and Review.

The wizard's model and method step with settings.

Watching a job

The job monitor.

The job monitor shows the stages and a live log. If a stage fails, the error is translated into a fix where possible, and Retry from checkpoint resumes after the stages that already finished.

Gate 3 · Scorecard

The scorecard.

The fine-tune against its prompt-only baseline on the held-out test set. Each metric says whether higher or lower is better, and changes are coloured accordingly. Samples show both answers side by side, with numbers the model invented flagged. A notice appears once a test set has been looked at: later runs judged on it are no longer a fair comparison.

Registry & playground

The registry.

Each version records its adapter, base model and every stage record. Convert → GGUF and Deploy to Ollama need a GPU host; without one they say so. The Playground answers from any model the stack serves, optionally side by side.

The playground.

2. A day in the life

An analyst on a pitch-writing team wants a small model that summarises startups the way the team writes — without inventing figures.

  1. 9:00 — Data. They upload 400 past pitch summaries. The report flags "$12.SM" in one row (a scanning error for $12.5M). They mark it Edit source, fix the source file, and upload it again. Everything is clean.
  2. 9:20 — Gate 1. Paging through the built examples, they notice one company makes up 60% of the rows; approval is blocked. They trim that company's examples, upload the file again, and build a new split. Approved.
  3. 9:35 — Gate 2. They pick a 0.5B Qwen model with QLoRA; Auto chooses Unsloth. Under Common they switch the scheduler to cosine and train for 2 epochs. The GPU check passes on the team's GPU box. They copy the equivalent command line into their notes, confirm, and start.
  4. 10:30 — Gate 3. The scorecard shows the task score up and invented numbers down against the prompt-only baseline. One sample still rounds a figure; they note it and ship — the registry now holds v1.
  5. 10:40 — Try it. After converting and deploying v1 on the GPU box, they compare it with the base model in the playground on a new pitch. The base model invents "40 hospital partners"; v1 doesn't.

Every decision above — who approved what, the notes, the exact settings — is recorded with the job.

3. Technical architecture

Components

ServiceRole
webnginx serving the single-page UI, proxying /api, sending security headers.
apiFastAPI control plane: projects, datasets, jobs and gates, runs, registry, serving proxy. CPU only.
workerarq worker running pipeline stages and registry tasks; streams logs to Redis.
postgresProjects, datasets, jobs, runs, reviews, registry versions (Alembic migrations).
redisJob queue and per-job log streams (served to the browser as server-sent events).
ollamaModel serving for the playground and deploys; models kept in a volume.
GPU hostPinned worker images (Hugging Face or Unsloth engine) reached through the compute backends (on-premises Docker; RunPod).

The pipeline

ingest → lint → build → split → [gate 1: data] → [gate 2: config] → train → evaluate → [gate 3: verdict] → package

Each stage reads the previous stage's output from the job's folder and writes its own, plus a manifest. A stage with a manifest is skipped on resume, so retries continue where they stopped. Gate decisions are manifests too. Approvals are atomic and locked, so two people approving at once can't run anything twice.

Parameters & validation

The list of training parameters is generated from the installed training libraries (trl's SFTConfig, peft's LoRA and VeRA configs, bitsandbytes, and Unsloth read from its source), with the library versions recorded. Parameters are grouped into trainer, peft and load, filtered by engine and method; secrets, multi-GPU options and settings the pipeline owns are never offered.

Settings are checked twice: at gate 2 against the list and a rules table (steps vs. epochs, one precision, evaluation needs a validation split, lengths within the model's context…), and again at the start of training by building the real library configs before any model is downloaded.

Reproducibility

Security

4. API reference

Base path /api behind the web front (the API itself listens on port 8000 on localhost). Every request needs the X-API-Key header. The full schema is served at /api/openapi.json.

EndpointPurpose
GET/POST /projectsList or create projects
POST /datasetsUpload a JSON Lines file (multipart: file, project_id)
GET /datasets, /datasets/{id}List or read datasets
GET /datasets/{id}/reportQuality report, lint, length bins, flagged numbers
POST /datasets/{id}/numbers/{row}Keep or edit-source decision on a flagged number
POST /datasets/{id}/build, /splitPreview built examples and a split
GET/POST /jobsList (filters: status, gate, project, dataset) or create a job
GET /jobs/{id}, /stages, /splitA job, its pipeline stages, its split
GET /jobs/{id}/logsLive log as server-sent events (resumable with Last-Event-ID)
POST /jobs/{id}/approve, /rejectDecide gates 1 and 2
POST /jobs/{id}/resplitReplace the job's split (at gate 1)
PUT /jobs/{id}/configSet the training configuration (at gate 2)
POST /jobs/{id}/dry-runTry settings on the CPU placeholder engine
POST /jobs/{id}/cancel, /retryCancel, or retry from the last checkpoint
GET /runs, /runs/{id}Runs (filters and sort) and a run's scorecard
POST /runs/{id}/verdictGate 3: ship, iterate or abandon
PUT /runs/baselinePin a project's baseline run
GET /runs/metricsEach metric and which direction is better
GET /registry, /registry/{id}Packaged versions
POST /registry/{id}/gguf, /deployQueue GGUF conversion or an Ollama deploy (GPU host)
GET /models, /engine, /preflightBase models, engine choice, GPU check
GET /train-paramsEvery parameter for a model, method and engine
GET /serve/models, POST /serve/{model}/generateServed models; a completion with latency and invented-number flags
GET /healthz, /metricsLiveness and job counts
curl -H "X-API-Key: $WHET_API_KEY" http://localhost:8080/api/jobs?status=awaiting_approval

5. Command line

export WHET_API_URL=http://localhost:8080/api WHET_API_KEY=<key>
whet project create demo
whet job create <project-id>
whet data report data.jsonl           # quality report, locally
whet params Qwen/Qwen2.5-0.5B-Instruct --method qlora [--all]
whet train Qwen/Qwen2.5-0.5B-Instruct data.jsonl --method qlora --max-steps 100 \
    --set trainer.lr_scheduler_type=cosine --set trainer.warmup_steps=0.1 \
    [--config train.yaml] [--eval-dataset val.jsonl] [--dry-run]
whet package ./run --version v1 --out ./registry
The wizard's review step shows the exact whet train command for the settings on screen, including every changed parameter as --set.

6. Configuration

VariablePurpose
WHET_API_KEYThe key the server accepts (required)
WHET_DATABASE_URL, WHET_REDIS_URLDatabase and queue (set by the stack)
WHET_HF_TOKENHugging Face token for gated models (server only)
WHET_MAX_UPLOAD_MBLargest upload (default 50)
WHET_OLLAMA_MODELModel the stack serves by default (default qwen2.5:0.5b)
WHET_OLLAMA_HOSTServing backend (default: the stack's Ollama)
WHET_HOST_ARTIFACT_DIRArtifacts folder on the Docker host, for GPU task containers

Local base models go in ./data/models, one folder per model with its config.json.

7. Known limits