Skip to main content

Y OS · on every Y Computer

Your AI, in one private workspace.

Plug it in and it already thinks. Frontier-class open models on silicon you own, the apps you chose wired to them, and nothing — prompts, files, history — ever leaving the building. That finished state is Y OS.

Introducing Y OS

Meet Y OS.

A custom software configuration, built for you as a service and shipped finished on every box. Scroll — we take the cover off and show you every layer.

Y OS architecture
DGX OSCUDA 13vLLMllama.cppTensorRT-LLMDeepSeek V4Qwen3.5gpt-oss-120bLaguna-SOpen WebUIQMOpenClawwhisperLiteLLM /v1TailscaleLUKSsnapshotsmanifest

The brains · local

Open models, running on the box.

The intelligence is local: open weights loaded onto your machine, picked by measured performance and pinned by hash. Everything else on this page connects to these.

Text & agentsMIT
DeepSeek V4 FlashDeepSeek · 304B A13BReads contracts, writes reports and reasons through hard problems without an external AI provider.Intelligence index 50 · 87GB on DS4 · 131K ctx
Long documentsApache 2.0
Qwen3.5 122BAlibaba · A10BReads a whole book in one go and answers questions about any page of it.Intelligence index 32 · 76GB · 262K ctx
CodeOpenMDW
Laguna-S-2.1Poolside · 118B A8BA senior programmer that writes, reviews and fixes code all day without a meter running.71GB NVFP4 · 31 tok/s measured
Text & agentsApache 2.0
gpt-oss-120bOpenAIOpenAI's open model — fast, familiar answers for everyday work.Intelligence index 24 · 59GB measured
Text & agentsMIT
Ling-3.0-flashinclusionAI (Ant) · 127B MoEThis week's frontier drop — big-model answers with half the box left over.IQ4_XS 66.4GB · NVFP4 81.4GB
FrontierOpen weights
GLM-5.2 (744B)Z.ai · 744BThe frontier itself — the biggest open model, pooled across a four-box cluster.405GB · vLLM TP=4 over RDMA
Text & agentsOpen weights
Mistral Medium 3.5Mistral AIThe reliable European generalist for writing, summarising and analysis.Intelligence index 30 · ~77GB at Q4
Speech-to-textMIT
Whisper large-v3OpenAITranscribes meetings, calls and voice notes — in almost any language.~3GB · faster than realtime

one authenticated endpoint · every app below runs on these brains

Ships your way

Order it with the apps you actually use.

Every app here thinks with the models above — one authenticated, OpenAI-compatible endpoint, no cloud in the loop. Each screenshot is the app's own, unmodified.

— and anything else that speaks the OpenAI API
Open WebUI — The default workspace

The default workspace. Order any box as-is and this is day one: team accounts, your pinned local models, documents with citations, no meter. Upstream's demo shown — yours boots wired to your local models.

Security baseline

Private by architecture. Boring by operation.

No mystery telemetry, no default public endpoint, no agent with the keys to the whole machine. The secure configuration is the default configuration.

  • 01

    Full-disk encryption and unique credentials during onboarding

  • 02

    Inference services bound to the machine, not exposed to the public internet

  • 03

    Remote access through a private tunnel—not an open port

  • 04

    Tools run with explicit permissions and isolated execution profiles

  • 05

    Model source, license, revision and file hashes recorded in a manifest

  • 06

    Diagnostics are customer-triggered and redact prompts and documents

  • 07

    Runtimes and models stay pinned until the combination passes our tests; updates snapshot first and can roll back

  • 08

    Web search, email and cloud models are off by default and labelled External — the data boundary is a setting you can see

Measured before handover

Your machine ships with its own numbers.

Every deployment ends with a benchmark certificate recorded on your exact unit: model hash, quantization, context, runtime, power profile, and single- and multi-request throughput. Not our marketing numbers — your machine’s.

Where our published numbers come from

Y OS benchmark certificate

Measured at handover

Y
Machine
NVIDIA DGX Spark · 128GB unified
Model
DeepSeek V4 Flash · UD-IQ3_XXS
Model hash
recorded at handover
Context window
131,072 tokens
Decode, single stream
16.2 tok/s
Prefill
411 tok/s
Power under load
~240W sustained

Specimen — your certificate records your exact unit

Y Recipes

See exactly what the machine can do.

Every recipe names the outcome, exact device, memory budget, software stack, measured limits and evidence grade. Follow the DIY path—or have Y OS build, secure, benchmark and maintain it.

Explore Y Recipes

01

Outcome

What useful work runs at the same time

02

Device

The exact memory, GPU and interconnect

03

Receipts

Raw results, sources and license status

04

Runbook

DIY steps and the Y OS-managed path

How you get Y OS

Installed on every Y Computer.

Mini 128, Pro 64 and Max 96 — one power user, parallel team services, or maximum CUDA compatibility. Order by card; the exact bill of materials is confirmed in writing before your machine ships.

From $5,999

founding-batch target · before tax and shipping

Compare Y Computers

Questions, answered plainly

Is Y OS a new Linux distribution?
No. It starts from NVIDIA's DGX OS and their own playbooks — we install the models, wire the apps you picked, pin every version, encrypt the disk, set up snapshots and benchmark the result. The value is the finished, supported machine, not software we pretend to have invented.
Can I buy Y OS without a Y Computer?
Not yet. For the founding release, Y OS ships installed on machines we configure, burn in and benchmark ourselves — that is how we can stand behind the result. A deployment offer for hardware you already own may come later.
Does it work without the internet?
Local chat, local document search, transcription and the local API can run without an external AI provider after the models are installed. Optional web search, email, cloud models and remote support require a network and are labeled External.
Is support a subscription?
The founding system price is one-time and includes the stated onboarding and first-year qualified update window. Any optional support after that will be quoted separately; local inference itself has no per-token charge.