Skip to main content

On its one task, a specialist model beats the API you’d deploy.

We put a model trained for one narrow job against the big-platform AI a team would reach for, score both on public data anyone can re-run, and publish all of it. The same test runs on your task.

What it costs to run

Daily queries:
Hosting scenario:
Model Monthly cost Annual cost

Self-hosted figures assume the GPU stays busy. Bursty or low volume narrows the gap, and below a volume threshold the hosted API wins on cost; the calculator is for finding your side of that line.

Every model and method, scored

Model Method Metric
Open-weight, fine-tuned (Apache 2.0)
Hosted API (cost tier)

Open any row’s Reproduce panel for its exact data, recipe, training health, cost, and verification hashes.

Where the errors go

Reproduce any number

Every result is reproducible from published artifacts. Run the released adapter on the eval data, or recompute the scores from the raw prediction logs. The base model is Apache-2.0 licensed, and every artifact is revision-pinned and hash-verified.

These are public-data results. See if it holds for your task.

Scope your task →

Prefer to talk? Book a free call →