A coding agent gives engineers leverage they can check — it writes code, the tests run, the diff is right there. Kith Bench is that, for analysts. It produces structured verdicts, extracted evidence and drafted deliverables under rules that make the output defensible: it never invents a number, nothing leaves as final until a named human approves it, and every run keeps a complete record of how it got there.
curl -fsSL https://bench.kithailab.com/install.sh | bash
The gap between "an LLM wrote something plausible" and "I'll sign my name under this" is not a prompting problem. It's a machinery problem. bench is that machinery: the model works inside a harness that constrains where facts come from, what counts as a finished answer, and who gets to call it final.
Chat is the front door — ask a question in plain language and bench routes it to the right specialist task. The audit trail, approval queue and deliverable export are the product, not developer tooling bolted on.
A pack is a folder: reasoning doctrine in versioned markdown, datasets, output schemas, vetted calculations, deliverable templates. No backend code. The flagship pack is climate risk assessment; the core is domain-agnostic.
Anthropic, Moonshot, OpenRouter — or point it at Ollama or LM Studio and run with no cloud provider and no egress at all. bench makes no network calls except to the providers you configure.
Every piece of work follows the same path — and each step leaves a record.
Not prompt instructions a model can drift away from — structural constraints that fail loudly when they're violated.
Values enter a run one of two ways: retrieving a dataset row, or invoking a pack-authored deterministic method — Python the agent can parameterise but never write, pinned by code hash and output hash. Anything cited that didn't come from one of those is a validation error, not a hallucination that slipped past a reviewer.
The approver is stamped from the login session. There is no API field to claim approval — you cannot automate away the signature.
Which model ran and why the router picked it (its verbatim reasoning, the candidates it considered), the doctrine version by content hash, every tool call, the token breakdown of the prompt itself, the dollar cost — and the estimated energy and carbon the run drew.
The reasoning rules live in versioned markdown the agent must follow and cite by heading. Changing how the analysis is done means editing a document a domain expert can read, not a prompt buried in code.
insufficient_data is a respectable verdict, and
missing data becomes an explicit data request rather than a guess. A run that
never lands a valid verdict ends
completed_without_output — a status that admits
there's nothing to read instead of implying there is.
On macOS or Linux, paste this into a terminal. It's the whole install.
curl -fsSL https://bench.kithailab.com/install.sh | bash
bash deserves a look first — read it here, it's short and commented./stop-bench.sh shuts it all down
bench opens at http://localhost:5180. Log in with
admin@example.com / bench-admin,
then add one AI provider key on the Settings page — a single OpenRouter key is
enough to try everything, and keys are stored encrypted, never in a file. Change
that password before anyone else can reach the machine.
| If you'd rather… | Do this |
|---|---|
| Not use a terminal | Download the repo as a ZIP and double-click start-bench.command (macOS) or start-bench.bat (Windows) — the plain-language walkthrough |
| Install nothing at all | Deploy to Render in one click — app, managed Postgres and a persistent disk, production secrets generated at deploy time (≈$13/mo) |
| Drive it yourself | git clone, then cp .env.example .env && docker compose up --build |
| Keep everything offline | Set BENCH_LOCAL_BASE_URL to your own inference server — zero-cloud setup |
A fresh install is already seeded with the climate-risk pack — sample sites, forward-looking regional signals and vendor-style hazard scores — so you can watch the whole machine work before setting up any of your own data.
"Is the vendor flood score for Alder Point still trustworthy?" The assistant recognises the request, delegates it to the right specialist harness task, and reports the draft verdict back — with the delegated run fully audited and its finding waiting in the approval queue.
Start a signal-divergence assessment from the Workbench and watch it retrieve data, reason under the doctrine and record a verdict — then read exactly what it cost, in dollars and in watt-hours.
Findings arrive as drafts. You read the evidence behind each one, and your approval is what turns it into something quotable.
Approved sections assemble into Markdown, HTML or a styled PDF — with an appendix listing the models used, the doctrine hash, approval status per section and the estimated footprint.
Every run carries an estimated energy and carbon number next to its dollar cost, and bench labels it an estimate everywhere it prints one — nothing here is metered. It's worth reporting because it's actionable: model choice moves it by more than an order of magnitude, and bench already picks the model. The full derivation, error bars included →
A harness-versus-unharnessed comparison — same model, same data, one arm inside bench and one with everything handed to it in context — is scaffolded in the open, with the case labels committed before any run so the git history proves they weren't fitted afterwards. No numbers are published until the full run is, limitations and all.
Installed packs are pinned by content hash and their methods are scanned for network access, process spawning and dynamic code — deterrents against accidents and drift. Installing a pack is still deploying code you reviewed.
Shipped credentials exist so one command works. Set
BENCH_ENVIRONMENT=production and bench refuses to
boot while the default secret key or admin password is still in place, rather
than serving insecurely.
Hardening checklist →