Open source · Apache-2.0 · runs on your own machine

AI knowledge work you can put in front of an auditor.

A coding agent gives engineers leverage they can check — it writes code, the tests run, the diff is right there. Kith Bench is that, for analysts. It produces structured verdicts, extracted evidence and drafted deliverables under rules that make the output defensible: it never invents a number, nothing leaves as final until a named human approves it, and every run keeps a complete record of how it got there.

Install — macOS & Linux needs Docker Desktop
$ curl -fsSL https://bench.kithailab.com/install.sh | bash
Windows, or no terminal at all? Other ways to run it →
Your models — Anthropic, OpenRouter, Moonshot, or fully local Your machine — no telemetry, ever Climate-risk pack included
What it is

A harness, not a chatbot

The gap between "an LLM wrote something plausible" and "I'll sign my name under this" is not a prompting problem. It's a machinery problem. bench is that machinery: the model works inside a harness that constrains where facts come from, what counts as a finished answer, and who gets to call it final.

Who it's for

Analysts, not engineers

Chat is the front door — ask a question in plain language and bench routes it to the right specialist task. The audit trail, approval queue and deliverable export are the product, not developer tooling bolted on.

What it knows

Pluggable domain packs

A pack is a folder: reasoning doctrine in versioned markdown, datasets, output schemas, vetted calculations, deliverable templates. No backend code. The flagship pack is climate risk assessment; the core is domain-agnostic.

Where it runs

Your models, your machine

Anthropic, Moonshot, OpenRouter — or point it at Ollama or LM Studio and run with no cloud provider and no egress at all. bench makes no network calls except to the providers you configure.

How it works

One run, end to end

Every piece of work follows the same path — and each step leaves a record.

01
You ask
In chat, or by starting a task directly from the workbench.
→
02
Harness loads
The task's doctrine, output schema and permitted tools — plus a model chosen for the job.
→
03
Facts retrieved
Dataset rows, or vetted calculations the agent can run but never rewrite.
→
04
Output checked
Every cited value is cross-checked against what was actually retrieved.
→
05
Human approves
A named person signs off; approved sections assemble into the deliverable.

The five rules, built into the architecture

Not prompt instructions a model can drift away from — structural constraints that fail loudly when they're violated.

  1. The AI never invents numbers

    Values enter a run one of two ways: retrieving a dataset row, or invoking a pack-authored deterministic method — Python the agent can parameterise but never write, pinned by code hash and output hash. Anything cited that didn't come from one of those is a validation error, not a hallucination that slipped past a reviewer.

  2. Outputs are drafts until a named human approves

    The approver is stamped from the login session. There is no API field to claim approval — you cannot automate away the signature.

  3. Every run is fully auditable

    Which model ran and why the router picked it (its verbatim reasoning, the candidates it considered), the doctrine version by content hash, every tool call, the token breakdown of the prompt itself, the dollar cost — and the estimated energy and carbon the run drew.

  4. Doctrine-as-context

    The reasoning rules live in versioned markdown the agent must follow and cite by heading. Changing how the analysis is done means editing a document a domain expert can read, not a prompt buried in code.

  5. Honest uncertainty

    insufficient_data is a respectable verdict, and missing data becomes an explicit data request rather than a guess. A run that never lands a valid verdict ends completed_without_output — a status that admits there's nothing to read instead of implying there is.

These guarantees are regression-tested. A golden-run suite drives the real engine against a scripted model, with each run locking in one of the five rules — so a reworded prompt or an edited doctrine file can't quietly stop them being true.
Install

Up and running in one command

On macOS or Linux, paste this into a terminal. It's the whole install.

$ curl -fsSL https://bench.kithailab.com/install.sh | bash

What the script does

  • Checks for Docker and git — and starts Docker Desktop if it's installed but asleep
  • Downloads bench into ~/kith-bench (re-run it later to update)
  • Creates the settings file for you — no config to edit
  • Builds and starts the containers, waits until bench answers, opens your browser
  • Prints your login and how to stop it again

Before you run it

  • You need Docker Desktop installed and opened once — the script tells you if it isn't
  • The first build takes several minutes and a few GB of disk; later starts take seconds
  • Piping a script to bash deserves a look first — read it here, it's short and commented
  • Everything lands in that one folder plus Docker's own storage; ./stop-bench.sh shuts it all down

After it starts

bench opens at http://localhost:5180. Log in with admin@example.com / bench-admin, then add one AI provider key on the Settings page — a single OpenRouter key is enough to try everything, and keys are stored encrypted, never in a file. Change that password before anyone else can reach the machine.

Other ways to run it

If you'd rather…Do this
Not use a terminal Download the repo as a ZIP and double-click start-bench.command (macOS) or start-bench.bat (Windows) — the plain-language walkthrough
Install nothing at all Deploy to Render in one click — app, managed Postgres and a persistent disk, production secrets generated at deploy time (≈$13/mo)
Drive it yourself git clone, then cp .env.example .env && docker compose up --build
Keep everything offline Set BENCH_LOCAL_BASE_URL to your own inference server — zero-cloud setup
First five minutes

It starts with a worked example, not an empty screen

A fresh install is already seeded with the climate-risk pack — sample sites, forward-looking regional signals and vendor-style hazard scores — so you can watch the whole machine work before setting up any of your own data.

Ask

Chat as the front door

"Is the vendor flood score for Alder Point still trustworthy?" The assistant recognises the request, delegates it to the right specialist harness task, and reports the draft verdict back — with the delegated run fully audited and its finding waiting in the approval queue.

Watch

A run you can follow

Start a signal-divergence assessment from the Workbench and watch it retrieve data, reason under the doctrine and record a verdict — then read exactly what it cost, in dollars and in watt-hours.

Approve

The queue is the control point

Findings arrive as drafts. You read the evidence behind each one, and your approval is what turns it into something quotable.

Ship

Deliverables with provenance

Approved sections assemble into Markdown, HTML or a styled PDF — with an appendix listing the models used, the doctrine hash, approval status per section and the estimated footprint.

Straight answers

What bench doesn't claim

The footprint figure is an estimate

Every run carries an estimated energy and carbon number next to its dollar cost, and bench labels it an estimate everywhere it prints one — nothing here is metered. It's worth reporting because it's actionable: model choice moves it by more than an order of magnitude, and bench already picks the model. The full derivation, error bars included →

Benchmark results aren't in yet

A harness-versus-unharnessed comparison — same model, same data, one arm inside bench and one with everything handed to it in context — is scaffolded in the open, with the case labels committed before any run so the git history proves they weren't fitted afterwards. No numbers are published until the full run is, limitations and all.

Pack integrity is a guardrail, not a sandbox

Installed packs are pinned by content hash and their methods are scanned for network access, process spawning and dynamic code — deterrents against accidents and drift. Installing a pack is still deploying code you reviewed.

The defaults are for your laptop

Shipped credentials exist so one command works. Set BENCH_ENVIRONMENT=production and bench refuses to boot while the default secret key or admin password is still in place, rather than serving insecurely. Hardening checklist →

Kith Bench is built by Kith AI Lab, where we do this work with clients. The platform is open source and free — Apache-2.0, no telemetry, no account required.