thinkingaloud

Asynchronous. Inference. Primitives.

I started this essay to introduce what I was building, but then I had to stop and go back to write two more (here and here) before publishing this. I had to capture the true intent and all the pivots in thinking, building and the evolving hypotheses behind what eventually turned out to be a set of asynchronous inference primitives.

In case this essay reached you first, here is the summary of the prequel.

So I was a new dad wanting to get the most out of AI but interleaved with changing diapers and bottle feeding. Chat sometimes was too much work for little flow in return.

The shape and size of my time were changing, and so was the work I was submitting to the models.

I was hoarding screenshots and wanted a way to synthesize them with AI and seamlessly surface and repurpose results whenever I needed. But small models and my ancient hardware made it hard to implement such workflows. Later when the ecosystem turned agentic the frameworks and API wrappers were too bundled. Powerful, but too big for my needs especially for working locally.

Then I figured out that most of the bundling was there to be compatible with the OpenAI-style v1/chat/completions REST API. So everything mostly was chat by default and other asynchronous-looking solutions also mostly were running a session with the model underneath. I found no alternatives to these synchronous primitive APIs.

So I pivoted from changing interfaces and managing screenshots to filling what I thought was a gap with asynchronous inference primitives.

Now back to this essay.

My first attempt at turning my abstract ideas into code and solving some personal pain points locally was called pnpl. Push Now and Pop Later (naming is hard even outside of Computer Science!). And the good thing about wanting to run local models on older hardware without a GPU was that I got introduced to llama.cpp! The impact of that open source runtime and its community is immeasurable. llama.cpp also continues to be the only constant through the changing phases of this side quest.

So when I came back the second time with the same hardware and different intent, the small model ecosystem was not small anymore.

The combination of Hugging Face and the llama.cpp community was pushing small models to the frontier!

Me and my hardware were set up for success with access to an ever-increasing number of the latest and multimodal models and the ability to run them seamlessly even on my old laptop.

For the first time I was getting work done with the primitives I was building while running models suited to my configuration.

And as they say, nothing works like work. Once I started to see results I could use, my center shifted.

I believe what you put in the center determines what gets built around it.

I started to see the current ecosystem as a galaxy with synchronous chat at its center. Over time, smaller systems formed within it around their own centers, for example coding, agentic sessions, model runtimes and enterprise AI.

Each such smaller system gives rise to its own orbit of tools, software, hardware, metrics, frameworks and economics, all conforming to its local center while still living within the larger synchronous chat based galaxy. So finding chat and chat based frameworks everywhere was a feature, not a bug!

All this led me to again change my hypothesis.

There are no gaps to fill, only centers to find and build around.

Once I stopped looking for gaps, it was easier to see why I could not find the primitives I wanted, and why even building them kept getting complicated. I was looking in the wrong orbit. Maybe even the wrong galaxy!

So roughly after three summers for me, during this second stint with building primitives, I tasted disproportionate flow for the first time. This newfound momentum and velocity was all because of my new center outside of the synchronous galaxy,

the work I wanted done asynchronously!

This recentering reinforced that when work is at the center, the intelligence does not have to be. It should still be powerful but served differently. Just the best generally available intelligence for the hardware served as a utility. So not all work has to be taken into a session by default to access it.

I did not want the lifetime of my work and its results to be tied to the runtime of a model.

Once intelligence became utilitarian, I could bring focus back to the center, my work. Ideas emerged about how I wanted to submit work as it arrives or in bulk without getting blocked. How I wanted the freedom to access results and repurpose them. Interesting questions emerged too,

But again I have been cautious not to end up bundling a queue service, a database, another server, timers and retries just to make the workflow seem asynchronous. I have been relying on the ideal stack already on my machine and everyone else's!

Files. Directories. Processes.

Filesystems have worked robustly for decades, letting the complexity, if any, build on top of them. They have stayed simple, swappable, small and incredibly useful.

So letting go of older orbits, recognizing the right center and looking closely at what always worked led me to find

nrvna.

A set of asynchronous inference primitives to make work and results durable.

nrvnad - run a model against a workspace

wrk - submit work to a workspace

flw - retrieve results from a workspace

I am cleaning up the repo and rewriting some docs, so you and your agents can build in this new galaxy centered around the work you want done asynchronously.