Nerdster Vault · Open-Model · nothing leaves the box
When your data can’t leave the building.
A dedicated open model runs entirely inside your own private, UK-hosted environment. No outside service sees your prompts, not us and not a model provider. For the firms whose rules leave no other option.
Book a free callStart with a fixed-price pilot from £2,500 (GPU hosting included), then from £1,200/mo single-tenant, or less on a shared private pool. Scoped in writing, no lock-in.
- Model runs in your own environment
- No third-party API, ever
- Air-gapped option
Why this exists
Some data can’t touch any outside model, even a private one.
The Gateway routes to managed models. This tier routes to nothing. The model lives on your own dedicated hardware.
For the strictest privileged, clinical or regulated work, even a zero-retention gateway is a step too far. Here we run an open model on a dedicated GPU inside your environment, so a prompt never crosses your boundary. It can even run fully air-gapped.
Models we run, and train
Open models we run for you, trained on your own data.
GPT-OSS-20B, our default
OpenAI’s open-weight 20B model: strong general reasoning at interactive speed on a single dedicated GPU. It’s efficient by design, so it stays fast without losing power.
Qwen3-Coder 30B and Qwen3 32B
Leading open models for code and deeper reasoning that still fit the same class of card. Our pick for code-heavy or analysis-heavy firms.
Llama 3.3, Gemma 3, Mistral
Proven alternatives we run where a firm prefers them or needs a bigger model, sized up to a larger GPU when the work calls for it.
Trained on your data, inside your walls
We can fine-tune the chosen model on your own documents, precedents and house style, so it answers in your voice and knows your work. Training runs inside your environment, so nothing leaves.
All open-weight, so there’s no vendor lock-in: when a better open model lands, we swap it in without changing how your firm works.
The honest limits
What a dedicated model can and can’t do.
Sized for a team, not a call-centre
A single dedicated GPU comfortably serves one busy team at over a hundred words a second. It isn't built to answer hundreds of people at once. That's the dedicated Intelligence Engine tier.
Long, not infinite, memory
The model can read a very long document, but a dedicated 20 GB card keeps roughly 30,000 words in view at once in practice. Bigger context needs a bigger GPU, which we can spec.
Excellent at your work, honest about frontier tasks
For private Q&A, drafting and first-pass review on your own data it's excellent. For the very hardest frontier reasoning, the Gateway’s top managed models still edge ahead, which is why many firms run both.
Right-sized, then grown
We start on the smallest GPU that does the job and move you up only when real usage calls for it, so you never pay for idle hardware.
FAQ
Vault Open-Model, answered.
How is this different from the Private Gateway?
The Gateway is private but calls managed models through a zero-retention layer. Open-Model calls nothing outside: a dedicated open model runs on your own GPU, inside your environment. It costs more because you have a GPU to yourself, but no prompt ever leaves your boundary.
Which model do you run, and why?
By default GPT-OSS-20B, OpenAI’s open-weight model, efficient enough to run fast on one dedicated GPU while giving strong general answers. For code-first firms we use Qwen3-Coder 30B. Because both are open-weight, we can upgrade you to a better open model later without disruption, and fine-tune whichever you choose on your own documents, entirely inside your environment.
Can it run air-gapped?
Yes. For the strictest cases the model and your data can run on infrastructure with no route to the public internet at all, so nothing can leave even in principle.
Can we share the cost of the GPU?
Where your compliance allows it, several clients can share one securely-isolated private GPU pool, which lowers the price. Where your rules require a GPU entirely to yourself, we run single-tenant at the higher figure.
What does it cost?
You start with a fixed-price pilot from £2,500, GPU hosting included, so you can judge a dedicated model on your own data first. If you keep it, from £1,200 a month single-tenant, or less on a shared private pool. The dedicated GPU is the main cost, and every figure is agreed in writing, with no lock-in.
The most private AI we build.
Book a free call. We'll size the right model and GPU for your work, show you the limits honestly, and agree the figures in writing before anything starts.
Book a free call0330 043 7414 hello@nerdster.ai Mon–Fri 9–5:30