Rendered at 19:17:03 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
neomantra 21 hours ago [-]
I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.
In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.
Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.
Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]
The project GitHub page is a much better introduction for the hn crowd.
Flere-Imsaho 6 hours ago [-]
I know HN folk don't like these kinds of websites with fancy animations and 80's style retro fonts...but I think this website actually contains the information that I needed in 1 page. I immediately knew what it was, what it required, and it provided the download links right there.
johnisgood 5 hours ago [-]
I just checked. "needed in 1 page"? What do you mean? Without scrolling through that mess?
3836293648 29 minutes ago [-]
1 page = 0 clicks, just scrolling. What you mean is called "above the fold"
LoganDark 6 hours ago [-]
I really don't think the animations or fonts are the problem. It's that nearly every LLM-generated landing page looks the same, and that they typically describe what the project is or what it does but not what it's like to use it. README usually contains what it looks like to use it, which is one of my biggest interests when seeing a project like this. I'm far less interested in SEO keyword-spam jargonslop in the tagline or title, even if it's descriptive and accurate!
For this project in particular I've known about ds4 for a while and the landing page feels like it's doing a gigantic disservice other than providing a download link. IMO, ds4 is far more interesting than this landing page would suggest!
simoiacos 23 hours ago [-]
Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
I can't run it because I don't have that hardware but that looks pretty neat, kudos!
aziis98 20 hours ago [-]
Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
simoiacos 20 hours ago [-]
Please open a PR! I was too conservative with the supported devices.
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
madduci 7 hours ago [-]
Can you share it? I have the same hardware, would be interested in trying it
simoiacos 5 hours ago [-]
I just added experimental support for Xe-LPG and Xe-LPG+ in main so you should be able to try it without a patch now.
As I said, I don't have the hardware to test it myself, so let me know how it goes!
madduci 4 hours ago [-]
Indeed! I let GLM5.3 write a patch as well, I will test it and submit to you as PR if it works. I've learned in the process that the 255H ships with DeepLink, which helps balancing the work between CPU and GPU (sounds like an OpenCL derivative?).
ilaksh 21 hours ago [-]
I wish someone would add Intel support to ds4. And also improve AMD support.
Maybe Intel and AMD should help them with that.
ABS 11 hours ago [-]
AMD sent antirez a Strix Halo back in June for this purpose
neomantra 4 hours ago [-]
There were ROCM commits to ds4 in late summer, so that's probably related. We include a ROCm build with ds4go (links elsewhere in this comment section), but it is absolutely untested by us whereas the Mac and DGX Spark are very tested.
simoiacos 21 hours ago [-]
Yeah I see the value but I built Xenolith to target smaller models.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
ilaksh 21 hours ago [-]
The recent Intel GPU/AI cards. Really the same type of models as ds4
ttoinou 21 hours ago [-]
Ive been using this since it was initially released with deepseek v4 flash, and it is absolutely the best launcher ever on my m5 max 128gb
Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?
haukebri 13 hours ago [-]
[flagged]
liuliu 21 hours ago [-]
If you are interested in high-end models with high-end Apple Silicon, also try out Local Code: https://releases.drawthings.ai/p/public-beta-of-local-code-b... It is currently in TestFlight (and will open-source next week), supporting vision with DeepSeek 4.1 Flash, Qwen 3.8 27B and DeepSeek 4 Flash 0731 without vision. Custom quants & SSD streaming to make these big models work with 64GiB and above devices (and of course, Qwen works with devices with 16GiB and above).
mark_l_watson 5 hours ago [-]
This looks excellent but I only have a 64G M5-Pro Mac so I probably need to stick with Sushi running Qwen3.8-Flash-Next.
quibono 7 hours ago [-]
What kind of models can you practically run with ds4 on a 64GB M5 Max?
> It must be more a working template for the biggest use cases, without trying to cover every possible setup.
this is such a great quote about building software in the AI age. given that everyone has a different use case, the best way for open source software to be built is to build it for the general use case and specialization could be done afterwards.
vlowther 23 hours ago [-]
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
darkwater 11 hours ago [-]
7800€ for a MacBook Pro M5 Max with 128GB RAM, right now. Insane.
trvz 5 hours ago [-]
Anyone who didn’t buy a 128GB Strix Halo system when they were 2500€ has only themselves to blame.
ttoinou 21 hours ago [-]
I’m already able to use 1M context windows with the same machine than you and same model. Strange
vlowther 20 hours ago [-]
Yeah, most of what I did was to add fused TQ support to leave more memory free for other nefarious purposes.
ttoinou 1 hours ago [-]
Oh that’d be nice indeed
wg0 15 hours ago [-]
Can someone explain me what expertise (domain knowledge) one needs to be able to write such model specific inference engine?
The other inference engine are also model by model with a huge switch statement deciding which part to load for which model or are they very generic?
epolanski 11 hours ago [-]
Antirez understood the fundamentals of how LLMs work, and read the inference code (or summaries of it via ai).
But what really made the difference was his understanding of hardware and systems programming in general and low level or architectural tricks to pull.
shieldagent 5 hours ago [-]
Local inference tooling from antirez is great news for small hardware. Same minimal-dependency philosophy that made Redis spread everywhere.
cuttothechase 18 hours ago [-]
Wondering how well this does with tool calling. Any one has any numbers or videos or anything using this?
From the github repo it seems like you really don't need a big Mac with huge amounts of RAM but SSD is sufficient.
If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!
neomantra 5 hours ago [-]
I just put up videos of some of the ds4go I mentioned. Most of it is long and boring, but it's there.
I started playing with local LLM+MCP in April 2025... I had to beg Qwen to look at the tool list and try anything.
These ds4 models, will happily call tools and all those harnesses I've made are composed of custom tools.
Once the HuggingFace+OpenAI showed how powerful notes are, I added a scratchpad tool to ds4go to improve self-improvement.
While you can do 64G/96G with the Qwen3.8 model, realistically you need 128G. Also, despite tons of playing with local models, the cloud-hosted models on bigger iron are smarter and faster. I don't truly code with my local models and don't recommend this path right now to replace something like Opus/Astra or full-brain DeepSeek4.
The "frontier-ness" of ds4 is great though! It has vast knowledge and thinking capability. Look at that steering video especially. I'm now exploring using ds4 for high-level thinking to create prompts for denser coding models.
gchamonlive 21 hours ago [-]
small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)
This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.
For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.
First of all, a x090 series card is "high end hardware" in its own right these days. Secondly, I'd think you'd probably get more interesting results running a MoE model in CPU-MoE mode, i.e. with the shared parameters residing on GPU and sparse experts on CPU plus SSD offload. Yes it will be slower, but small dense models are just a dead end and not that interesting. (Note that prefill would still be sped up in this setting; the CPU/GPU layer split in llama.cpp and the like applies to decode, but even a "0 graphic layers" setup does accelerate prefill.)
jminnl 6 hours ago [-]
3090 and 4090 are amazing value for the money on the second hand market. 5090's are sold for more than they were worth when they were new and new ones are ridiculously priced.
gchamonlive 20 hours ago [-]
Qwen3.8 35ba3b is decidedly faster but also dumber, at least in my tests.
latentsea 15 hours ago [-]
> Qwen3.8 35ba3b
Does not exist. You're thinking of Qwen3.6 35ba3b
TheqO 10 hours ago [-]
Metal, the primary target, on Macs with 96 GB or more. Smaller machines can use SSD streaming. SSD streaming is also needed in order to run very large models such as full GLM 5.x (not Flash) on 128GB systems
Anyone tested token speeds at less than 96gb RAM on apple?
HoldOnAMinute 22 hours ago [-]
How is this different from other LLM runners?
simonw 21 hours ago [-]
It's more likely to work. Most LLM runners are meant to work with any model, which means there are all kinds of ways you might misconfigure them in a way that causes function tooling not to work, or performance to be less than you would like.
DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.
locknitpicker 14 hours ago [-]
I'm sorry, it's hard for me to understand what point you were trying to make. So existing LLM runners are designed to support all models, and they run all models, but they might be misconfigured? And DS4 is better because it's unable to run all models?
simonw 13 hours ago [-]
It has better defaults.
aflinik 9 hours ago [-]
Does that imply that with enough effort spent on fine-tuning the configuration, other LLM runners can achieve similar results?
rrgok 10 hours ago [-]
So wouldn't be easier just to provide a repository of better defaults for each model? Like lsp-config for neovim?
jminnl 6 hours ago [-]
The problem is that with diverse hardware such a set of defaults is much harder to make. You get this matrix of possibilities: gpu, VRAM and memory configurations and then the model axis. This leads to way too many options. The better alternative would be to have the runner self-benchmark what the best settings are given that it already has access to that one particular configuration.
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
llama-bench is next to useless for this purpose.
ilaksh 21 hours ago [-]
Emphasis on performance and usable coding/agentic ability for consumer AI hardware. Does not attempt to handle all models or hardware at once but rather focuses on optimizing the best options for that category of hardware.
ttoinou 19 hours ago [-]
Lots of small details are taken care of so it runs smoothly. For example ds4-agent is append only, never rewriting history of messages, keeping KV cache prefix reusable. Huge benefit
pydry 21 hours ago [-]
My instinctive reaction from the readme is that it isnt. It's apparently a vibe coded knock off of llama.CPP.
jminnl 6 hours ago [-]
The llama.cpp guys don't get nearly enough credit for their work. Though the quality of the codebase is dropping over time, it is still quite high compared to most of the alternatives, and it is still one of the most stable ways to run a large variety of models.
Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).
csmlab_notes 21 hours ago [-]
[flagged]
mannyv 15 hours ago [-]
Engineers have entered the building. We've come a long way from people debating whether mmap was safe to use.
wg0 15 hours ago [-]
If creator of Redis is usig AI to write serious software, ordinary folks need to rethink their stance.
Flere-Imsaho 6 hours ago [-]
I'm still meeting software people who are still very anti-AI...this is despite the strong evidence that a frontier model is superior to 99% of software engineers at writing code, producing documentation, testing, generating threat models, etc. IF prompted correctly.
As Antirez is using a non-frontier model for his work (via locally-running), then I think that is further proof that AI is ready for widespread use in software engineering.
jminnl 6 hours ago [-]
AI can produce code faster than you can review it and if you are not careful you find yourself on the other side of a trapdoor with absolutely no way back. Your codebase has become a mess that you can only maintain with more AI. But if you are careful and prune regularly you can do well and gain a very good increase in development speed.
The risks are:
- creating a lot of dead code
- ending up with substantial repetitions (this is getting better over time but the risk is definitely still there)
- testing only on the happy path rather than all execution paths
- mixing current and outdated information resulting in subtly broken code
- inability to reason past a certain level of complexity, but no signal that this is the case
For each of these risks there are remedies, one of the more powerful ones for me is the ability to just roll back when things have gone too far off the rails, realize in hindsight what caused it and to retry with a much better initial prompt.
dudefeliciano 9 hours ago [-]
The creator of this has a very interesting YouTube channel he posts to almost daily, talking mostly about current developments on AI from a technical but also societal/philosophical point of view. I specifically like that he provides a (much needed in this space) leftist point of view while not being anti-AI. Most of the videos are in Italian, so if you speak Italian (or are fine with YouTube automatic translation) I highly recommend it:
His blog is good as well, but he doesn't post there as frequently: http://antirez.com
yieldcrv 17 hours ago [-]
I’m a little confused
ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time
but its model agnostic-ish
and benchmarks compared to what? what do these large MoE models typically get in tokens per second?
I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?
I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
futhey 5 hours ago [-]
I could be wrong but my read from following this on Twitter was that it started on Deepseek 4 (flash). It probably started out focused solely on that model with no guarantee the techniques would transfer to anything else.
Almondsetat 20 hours ago [-]
This website is pure slop. I'd ask @dang to just link the original repo
aeve890 17 hours ago [-]
Right? Compare this with antirez's blog lmao. The very author of an incredible piece of software using the most plain website possible, while a derivative post about the same tool it's a slop fest with useless FX, cringe hackerman style palette and such. It's just too funny.
timmytokyo 15 hours ago [-]
Here's a sample of the site's headers. Note the heavy reliance on slop marketing-speak (rule of 3, X not Y, etc.).
"Compressed, not lobotomized."
"Dense, resident, yours."
"Local frontier inference, narrow on purpose."
"ds4 hardware fit: local, streamed and distributed."
The whole site says nothing with so many words. It's also got all the hallmarks of a typical vibe-coded web site (small all-caps text, highly sectioned content, silly animations). Why do people do this? It doesn't impress. In a few years, we'll look back on sites like this like we look at geocities sites today.
aeve890 15 hours ago [-]
>like we look at geocities sites today.
We look at geocities with nostalgia, I guess. Ugly as fuck but made with heart when all this thing of the internet was growing.
This slop shit on the other hand... It's cringe right now.
pulkitsh1234 22 hours ago [-]
curious, why did antirez go with C instead of something like Rust ?
ilaksh 21 hours ago [-]
Antirez has been writing C for a million years so is much more familiar with it than Rust.
Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).
Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?
Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.
zozbot234 20 hours ago [-]
> Antirez has been writing C for a million years so is much more familiar with it than Rust.
This is explicitly an AI-coded project, Antirez argues that LLMs are worse at writing Rust than C because so much high quality systems code (think e.g. sendmail) that ends up in AI training sets is C, not Rust. Another related argument is that the more detailed syntax and compiler feedback found in Rust compared to C are really a negative for LLM workflows.
There's plenty of room to disagree wrt. this of course: without the strong typing checks of Rust around e.g. indirect references, safety and correctness ends up being a global property in typical C programs, and LLMs are terrible wrt. reasoning about global properties. You're better off forcing them to adapt to a different local syntax that does a more complete job of enforcing modularity, since this is comparatively foolproof.
cuttothechase 18 hours ago [-]
Yes, it is AI-coded. But definitely not a one shot kind of a deal.
Much easier to work with a language you are most comfortable with right?
Aeolos 21 hours ago [-]
Yes, Rust gives you great access to low-level code on different platforms, including SIMD. It is also alias-free by default, and gives you excellent primitives to write multi-threaded code with compile-time correctness guarantees, which is how projects such as zlib-rs end up significantly faster than their C counterparts.[1]
It's about as good as it can get for this kind of code.
He recently said that he finds Rust less ergonomic and that this also affects code written by LLMs, which he thinks excel at writing C partly because of the enormous, high-quality codebase they were trained on. He sees security-critical code as a reason to choose Rust.
Personal preference of the author, he made at least one video on YouTube on why he dislikes Rust. I think he finds it too cumbersome and not worth it when the software isn't security-critical (not that I agree, just reporting what IIRC his stance is).
wg0 15 hours ago [-]
Because C is the simplest language that a competent programmer learn just in an afternoon pretty much.
I like C's simplicity so much. The only other language that comes close in simplicity and minimalism is go.
Flere-Imsaho 6 hours ago [-]
So if LLMs can write C really "well" - as in they can keep track of all the memory allocations, branches and conditions that would prevent the typical memory problems associated with C...then do we even need Rust anymore? The control of the memory allocations and layout in C does theoretically mean you can ultra-optimise the code. The LLMs can write 1000s of unit tests and they're really good at fuzzing.
I don't know Rust well enough to understand what else it would provide over the safe memory guarentees?
zozbot234 4 hours ago [-]
> So if LLMs can write C really "well" - as in they can keep track of all the memory allocations, branches and conditions that would prevent the typical memory problems associated with C...then do we even need Rust anymore? The control of the memory allocations and layout in C does theoretically mean you can ultra-optimise the code. The LLMs can write 1000s of unit tests and they're really good at fuzzing.
If.
(Mind you, fuzzing a program with any non-trivial input space can only ever prove that it is unsound. It can never prove that the program is sound.)
xlayn 16 hours ago [-]
In case you like the store kv to disk so you can resume I keep this branch of llama.cpp that includes that same functionality
And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...
jasonjmcghee 16 hours ago [-]
Last time I was using llama cpp you could just do:
llama_state_save_file
or
llama_state_seq_save_file
and the load equivalents.
That was a year or so ago though...
try-working 21 hours ago [-]
There are insane speed improvements for local inference going around on X right now. They've popped up the last month and week.
Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week.
There's lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc.
I'm really excited for this as I'll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it's a 96gb machine), and they may come close in performance to DS 4/4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.
locknitpicker 14 hours ago [-]
What does ds4 offer that projects such as llamma.cpp or ollama haven't been offering for a while?
I mean, I've been using local models on vscode right next to frontier models with ollama for a few months. What's new?
doctorpangloss 1 days ago [-]
the problem is the dsv4 checkpoint so quantized isn't very good
ilaksh 22 hours ago [-]
Which ds4 checkpoint for which model exactly did you test? Don't they have multiple different versions and quantization levels?
He had an awesome opportunity to do high concurrency synchronous replication (raft) on top of in-memory databases at a time where ssds were still uncommon, but instead chose to redneck-engineer his own protocol, then double-down that he knows best.
Not that dissimilar to choosing C over rust for familiarity.
tuesdaynight 9 hours ago [-]
>Not that dissimilar to choosing C over rust for familiarity.
It's LLM age. Just port it to Rust if that is so wrong for you. He is doing it for free, no need for arguing about the language he wants to use
za_creature 1 hours ago [-]
I'm... not arguing?
Dude asked what's wrong with antirez, I answered
jacquesm 4 minutes ago [-]
I don't see much wrong with antirez. He does what he does and gives it to the world for free, if you don't like it you are free to roll your own or ignore it. All you are doing here is behaving in a jealous way that makes no sense to me.
elktown 10 hours ago [-]
That quarrel was so incredibly petty. Oneupmanship, madness of not conforming with the zeitgeist, cherry-picking, whatnot.
Yeah, as usual with devs; pick a tribe then go to insufferable lengths with the newfound and completely unearned superiority complex.
za_creature 10 hours ago [-]
aphyr: distributed systems are difficult and break in ways that are difficult to predict, this is known scientific fact and here's a long list of databases I broke because their engineers think the rules don't apply to them.
antirez: no, u!
elktown: both sides are tribals with superiority complexes!
elktown 9 hours ago [-]
I meant you.
za_creature 58 minutes ago [-]
Did I insult your god and savior by stating that he fucked up?
Cause he did fuck up, then changed the subject, then claimed the criticism was unfair because he was clearly talking about something else.
In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.
Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.
Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]
EDIT: add ds4go TUI screenshot gist [4]
[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...
[2] https://github.com/nimblemarkets/ds4go#install
[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...
[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...
The project GitHub page is a much better introduction for the hn crowd.
For this project in particular I've known about ds4 for a while and the landing page feels like it's doing a gigantic disservice other than providing a download link. IMO, ds4 is far more interesting than this landing page would suggest!
I'm also looking into expanding the protocol and the engine to support various steering techniques.
https://github.com/simoneiacomino/xenolith
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
As I said, I don't have the hardware to test it myself, so let me know how it goes!
Maybe Intel and AMD should help them with that.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?
> It must be more a working template for the biggest use cases, without trying to cover every possible setup.
this is such a great quote about building software in the AI age. given that everyone has a different use case, the best way for open source software to be built is to build it for the general use case and specialization could be done afterwards.
The other inference engine are also model by model with a huge switch statement deciding which part to load for which model or are they very generic?
But what really made the difference was his understanding of hardware and systems programming in general and low level or architectural tricks to pull.
From the github repo it seems like you really don't need a big Mac with huge amounts of RAM but SSD is sufficient.
If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!
https://gist.github.com/neomantra/d49df05d6b137b9e6844186499...
I started playing with local LLM+MCP in April 2025... I had to beg Qwen to look at the tool list and try anything.
These ds4 models, will happily call tools and all those harnesses I've made are composed of custom tools.
Once the HuggingFace+OpenAI showed how powerful notes are, I added a scratchpad tool to ds4go to improve self-improvement.
While you can do 64G/96G with the Qwen3.8 model, realistically you need 128G. Also, despite tons of playing with local models, the cloud-hosted models on bigger iron are smarter and faster. I don't truly code with my local models and don't recommend this path right now to replace something like Opus/Astra or full-brain DeepSeek4.
The "frontier-ness" of ds4 is great though! It has vast knowledge and thinking capability. Look at that steering video especially. I'm now exploring using ds4 for high-level thinking to create prompts for denser coding models.
For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.
I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server
Does not exist. You're thinking of Qwen3.6 35ba3b
DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
llama-bench is next to useless for this purpose.
Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).
As Antirez is using a non-frontier model for his work (via locally-running), then I think that is further proof that AI is ready for widespread use in software engineering.
The risks are:
- creating a lot of dead code - ending up with substantial repetitions (this is getting better over time but the risk is definitely still there) - testing only on the happy path rather than all execution paths - mixing current and outdated information resulting in subtly broken code - inability to reason past a certain level of complexity, but no signal that this is the case
For each of these risks there are remedies, one of the more powerful ones for me is the ability to just roll back when things have gone too far off the rails, realize in hindsight what caused it and to retry with a much better initial prompt.
https://youtube.com/@antirez
ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time
but its model agnostic-ish
and benchmarks compared to what? what do these large MoE models typically get in tokens per second?
I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?
I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
"Compressed, not lobotomized."
"Dense, resident, yours."
"Local frontier inference, narrow on purpose."
"ds4 hardware fit: local, streamed and distributed."
The whole site says nothing with so many words. It's also got all the hallmarks of a typical vibe-coded web site (small all-caps text, highly sectioned content, silly animations). Why do people do this? It doesn't impress. In a few years, we'll look back on sites like this like we look at geocities sites today.
We look at geocities with nostalgia, I guess. Ugly as fuck but made with heart when all this thing of the internet was growing.
This slop shit on the other hand... It's cringe right now.
Also the goal of the project is to squeeze the absolute maximum performance and capability possible out of limited hardware resources (compared to clusters of B200s or something).
Does Rust even give you good access to low-level code on different platforms? And if so, how much extra work do you need to do to make it acceptable to the compiler? And is that work worthwhile if you are not going to get the security guarantees of normal Rust code? Is it a worthwhile tradeoff when the goal is performance?
Those are real questions by the way, not rhetorical. If Rust could work well for this type of project then I would like to know.
This is explicitly an AI-coded project, Antirez argues that LLMs are worse at writing Rust than C because so much high quality systems code (think e.g. sendmail) that ends up in AI training sets is C, not Rust. Another related argument is that the more detailed syntax and compiler feedback found in Rust compared to C are really a negative for LLM workflows.
There's plenty of room to disagree wrt. this of course: without the strong typing checks of Rust around e.g. indirect references, safety and correctness ends up being a global property in typical C programs, and LLMs are terrible wrt. reasoning about global properties. You're better off forcing them to adapt to a different local syntax that does a more complete job of enforcing modularity, since this is comparatively foolproof.
Much easier to work with a language you are most comfortable with right?
It's about as good as it can get for this kind of code.
[1] https://www.reddit.com/r/rust/comments/1ixt1ei/zlibrs_is_fas...
The video is in Italian but has an auto-dubbed English audio track: https://www.youtube.com/watch?v=sOt0WpQG5eU\&t=526s
I like C's simplicity so much. The only other language that comes close in simplicity and minimalism is go.
I don't know Rust well enough to understand what else it would provide over the safe memory guarentees?
If.
(Mind you, fuzzing a program with any non-trivial input space can only ever prove that it is unsound. It can never prove that the program is sound.)
https://github.com/alainnothere/llama.cpp/commits/disk-cache...
And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...
That was a year or so ago though...
Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week.
There's lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc.
I'm really excited for this as I'll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it's a 96gb machine), and they may come close in performance to DS 4/4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.
I mean, I've been using local models on vscode right next to frontier models with ollama for a few months. What's new?
the ds4 quants were very good beating the unsloth quants https://github.com/michaelasper/benchmarks/blob/main/deepsee...
What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?
https://aphyr.com/posts/283-jepsen-redis
https://antirez.com/news/55
finally
https://aphyr.com/posts/307-jepsen-redis-redux (see his comments there too)
He had an awesome opportunity to do high concurrency synchronous replication (raft) on top of in-memory databases at a time where ssds were still uncommon, but instead chose to redneck-engineer his own protocol, then double-down that he knows best.
Not that dissimilar to choosing C over rust for familiarity.
It's LLM age. Just port it to Rust if that is so wrong for you. He is doing it for free, no need for arguing about the language he wants to use
Dude asked what's wrong with antirez, I answered
Yeah, as usual with devs; pick a tribe then go to insufferable lengths with the newfound and completely unearned superiority complex.
antirez: no, u!
elktown: both sides are tribals with superiority complexes!
Cause he did fuck up, then changed the subject, then claimed the criticism was unfair because he was clearly talking about something else.