Skip to main content

Y Recipes · Open systems library

Run the impossible.Own the result.

Exact models. Exact machines. Exact numbers. Y Recipes turns frontier AI into private systems you can inspect, reproduce and run every day.

Measured by YSame machine · same weights · same prompt

284B.One DGX Spark.

Y runs DeepSeek V4 Flash 0731 as a private, OpenAI-compatible endpoint at 28.29 tokens/sec 1.67× the target-only path on the exact same desk-side machine.

Flagship proof 001Passed

28.29 tok/s

fixed generation

1.67×

vs target-only

9.07 GiB

memory left

6/6

integration checks

32K

configured context

18K

longest tested input

0 KiB

process swap

Time to answer

Mean wall time for the same fixed 256-token response · lower is faster

15.41 s
Target
9.40 s
Upstream
9.34 s
Y

Speed of intelligence

Generated tokens per second · higher is faster

16.93 t/s
Target
28.09 t/s
Upstream
28.29 t/s
Y

Y optimization layer

Same speed. Smaller sidecar. More room to work.

We compressed the speculative sidecar by 21.64% while keeping measured throughput level with upstream and recovering 2.06 GiB of observed memory headroom.

Exact DGX Spark runUpstreamY
Speculative sidecar10.15 GiB7.95 GiB21.64% smaller
Minimum memory available7.01 GiB9.07 GiB+2.06 GiB
Integration smoke6/66/6same gate

Method. Same DGX Spark, target weights, runtime, 32K configuration, prompt, 256-token output, seed and five measured runs per profile. One warm-up excluded. This is a fit, speed, memory and integration result; intelligence quality is reported separately on a common evaluation harness.

Reproduce it

What Y runs

From one model to an entire company.

One person can own the same intelligence stack that used to require an infrastructure team. Y OS installs it, tunes it and keeps it private.

01

A private frontier reasoning stack

DeepSeek V4 Flash, 284.3B parameters, one OpenAI-compatible endpoint on one DGX Spark.

28.29 tok/s · 32K configured · 0 KiB swap

Build the DeepSeek stack
02

Agents that make finished media

Connect a long-context reasoning model to image, video, voice and automation tools behind one private network.

Agent endpoint · video lanes · workflow automation

See the video factory
03

Local models on phones and laptops

Use the same Y model catalog to choose the smallest useful quant for a phone, laptop, Mini, Pro or Pro Max.

Device-native models · no forced cloud dependency

Explore the model catalog

Built with open AI

This is what one person can make.

Image, video, voice, coding, research and automation—one private intelligence stack, connected to the tools you already use. Every card links to the model or open-source interface behind it.

A matte-black mini computer on a walnut desk, generated with Qwen-Image-2512Output sample

Image generation

Qwen-Image-2512

Generate product concepts, campaign art and visual worlds with an open image model connected to your private workspace.

Y measurement: 63GB working set · 212s per image on one Spark.

Output sample

Video generation

Wan 2.2

Turn a written shot into finished video with an open text-to-video model inside the same private stack.

Codex CLI coding agent interfaceOpen-source UI

Coding agents

Codex CLI

Point compatible coding agents at a private OpenAI-compatible endpoint and build without sending the repository to a model host.

Open WebUI private AI workspace interfaceOpen-source UI

Private research

Open WebUI

A polished private workspace for chat, research, documents and model routing across the Y OS model library.

Output sample

Voice generation

Chatterbox

Create private narration, character voices and product audio with the published MIT-licensed model.

n8n automation workflow interfaceOpen-source UI

Automations

n8n

The upstream workflow builder. Its AI nodes can call the same private, OpenAI-compatible endpoint as every other tool.

Ideas worth running

What people are building right now.

We track the strongest new releases and workflows from X, GitHub and model labs, then turn the useful ones into repeatable Y Recipes.

August 2026 model radar

The latest models. Sized for real devices.

Phone, laptop, desk-side or server: we track accessible weights, commercial terms, runtime support and the smallest credible Y system for each release.

Liquid AI · August 2026

Ready · phone class

LFM2.5-2.6B

This one is actually phone-class.

Official GGUF quants start at 1.59GB. Liquid reports 30 tok/s on a phone and under 2.5GB of memory for its on-device agent.

Syzygy Research · August 2026

Downloadable beta · laptop class

Mach-1 Additive 35B

Impressive compression. Not a phone model.

The public Apache-2.0 checkpoint is about 8GB and reconstructs a Qwen3.6-35B-A3B teacher using additive weights.

Pokee AI · August 2026

Gated · no independent proof

Pokee-Isaac 28B

A launch post is not a shippable model.

Pokee advertises a 10M-token agent and single-GPU deployment, but the 28B checkpoint is not publicly downloadable for an anonymous audit.

DeepSeek · July 31, 2026

Y measured · Spark class

DeepSeek V4 Flash 0731

284.3B parameters. One DGX Spark.

Y loaded the 97.05 GiB target, served it at a fixed 32K context and reached 28.29 tok/s with the Y IQ3_M speculative sidecar. The unchanged six-case integration gate passed 6/6 with 9.07 GiB of minimum observed memory headroom.

Built in public

Fork the recipe. Inspect the proof. Or let Y do it all.

The public repository contains the pinned build, hashes, commands, raw measurements and machine-readable results. Y OS turns that work into a supported appliance.