A private frontier reasoning stack
DeepSeek V4 Flash, 284.3B parameters, one OpenAI-compatible endpoint on one DGX Spark.
28.29 tok/s · 32K configured · 0 KiB swap
Build the DeepSeek stackY Recipes · Open systems library
Exact models. Exact machines. Exact numbers. Y Recipes turns frontier AI into private systems you can inspect, reproduce and run every day.
Y runs DeepSeek V4 Flash 0731 as a private, OpenAI-compatible endpoint at 28.29 tokens/sec— 1.67× the target-only path on the exact same desk-side machine.
28.29 tok/s
fixed generation
1.67×
vs target-only
9.07 GiB
memory left
6/6
integration checks
32K
configured context
18K
longest tested input
0 KiB
process swap
Time to answer
Mean wall time for the same fixed 256-token response · lower is faster
Speed of intelligence
Generated tokens per second · higher is faster
Y optimization layer
We compressed the speculative sidecar by 21.64% while keeping measured throughput level with upstream and recovering 2.06 GiB of observed memory headroom.
Method. Same DGX Spark, target weights, runtime, 32K configuration, prompt, 256-token output, seed and five measured runs per profile. One warm-up excluded. This is a fit, speed, memory and integration result; intelligence quality is reported separately on a common evaluation harness.
Reproduce itWhat Y runs
One person can own the same intelligence stack that used to require an infrastructure team. Y OS installs it, tunes it and keeps it private.
DeepSeek V4 Flash, 284.3B parameters, one OpenAI-compatible endpoint on one DGX Spark.
28.29 tok/s · 32K configured · 0 KiB swap
Build the DeepSeek stackConnect a long-context reasoning model to image, video, voice and automation tools behind one private network.
Agent endpoint · video lanes · workflow automation
See the video factoryUse the same Y model catalog to choose the smallest useful quant for a phone, laptop, Mini, Pro or Pro Max.
Device-native models · no forced cloud dependency
Explore the model catalogBuilt with open AI
Image, video, voice, coding, research and automation—one private intelligence stack, connected to the tools you already use. Every card links to the model or open-source interface behind it.
Output sampleImage generation
Video generation
Open-source UICoding agents
Open-source UIPrivate research
Voice generation
Open-source UIAutomations
Ideas worth running
We track the strongest new releases and workflows from X, GitHub and model labs, then turn the useful ones into repeatable Y Recipes.
From X · Unsloth
A useful open model running entirely on a phone—the smallest edge of the Y model library.
Phone demo · Google
Google's official Android gallery shows what local multimodal AI already feels like in your hand.
Agent demo · Microsoft
Microsoft's computer-use agent demonstrates workflows a private Y endpoint can orchestrate.
Official use cases · Qwen
Qwen's official examples show the coding, tool-use and repository workflows local models can drive.
August 2026 model radar
Phone, laptop, desk-side or server: we track accessible weights, commercial terms, runtime support and the smallest credible Y system for each release.
Liquid AI · August 2026
Ready · phone classThis one is actually phone-class.
Official GGUF quants start at 1.59GB. Liquid reports 30 tok/s on a phone and under 2.5GB of memory for its on-device agent.
Syzygy Research · August 2026
Downloadable beta · laptop classImpressive compression. Not a phone model.
The public Apache-2.0 checkpoint is about 8GB and reconstructs a Qwen3.6-35B-A3B teacher using additive weights.
Pokee AI · August 2026
Gated · no independent proofA launch post is not a shippable model.
Pokee advertises a 10M-token agent and single-GPU deployment, but the 28B checkpoint is not publicly downloadable for an anonymous audit.
DeepSeek · July 31, 2026
Y measured · Spark class284.3B parameters. One DGX Spark.
Y loaded the 97.05 GiB target, served it at a fixed 32K context and reached 28.29 tok/s with the Y IQ3_M speculative sidecar. The unchanged six-case integration gate passed 6/6 with 9.07 GiB of minimum observed memory headroom.
Built in public
The public repository contains the pinned build, hashes, commands, raw measurements and machine-readable results. Y OS turns that work into a supported appliance.