The Stack

Software, APIs, and things you can use

Open Source Radar — September 24: The agent toolchain hardens

The Stack

Open Source Radar — September 24: The agent toolchain hardens

Today's trending page has almost nothing to do with new models. It is all scaffolding — harnesses, code intelligence, and the plumbing that makes ordinary software callable by an agent. impeccable (JavaScript, ~70,500 stars) A design language for AI coding agents, and a direct attack on how AI-built frontends look. It ships one skill, 23 commands, and 61 deterministic detector rules aimed at the standard tells: the same geometric sans-serif for everything, purple-to-blue gradients, cards neste

Open Source Radar — September 23: the agent's missing plumbing

The Stack

Open Source Radar — September 23: the agent's missing plumbing

Today's board is the three layers an agent needs before anyone trusts it with real work: a document surface it can actually edit, a runtime that keeps it isolated and cheap to wake, and a way to hand it tools without handing it keys. Univer (TypeScript, ~16,000 stars, Apache-2.0) — An office SDK that has spent three years being the open-source spreadsheet you embed in your own product, and has now repositioned itself, in its own README, as the Office Harness for AI Agents. Spreadsheets, documen

AI 101 — What is a TPU?

The Stack

AI 101 — What is a TPU?

A TPU — Tensor Processing Unit — is a computer chip Google built to do one thing: the enormous matrix multiplications that neural networks run on, done faster and with less electricity than a general-purpose chip can manage. It is the silicon underneath Search, Translate and Gemini, it is what Google rents to other AI companies through its cloud, and in 2026 it became collateral in some of the largest loans the AI buildout has seen. If you have read about a "$22 billion loan" or "a million TPUs"

Rabbit's OS3 agent runs on your laptop — and the R1 is done

The Stack

Rabbit's OS3 agent runs on your laptop — and the R1 is done

Rabbit spent two years as the cautionary tale of AI hardware. On Tuesday it started selling the software instead: an agent that runs on the machines you already own. Rabbit has launched OS3, a standalone "agentic operating system" that lives in the cloud but executes work through a local agent node you install on Windows, Mac, or Linux — and founder Jesse Lyu told Wired the company has stopped manufacturing the R1, the pocket gadget that made it famous. The R1 was not, by Lyu's account, a comme

Google's Intrinsic open-sources its robot control stack

The Stack

Google's Intrinsic open-sources its robot control stack

Two physical-AI stories landed on the same day, and together they draw the line the industry is converging on: the software layer is going free, and the sensing layer is going into the hands of one buyer. Alphabet's Intrinsic open-sourced Intrinsic Core, the runtime, SDK and real-time control framework it had been keeping in-house — an Apache-2.0 layer that handles trajectory adjustment mid-movement, 6-DoF pose estimation, collision-avoidance motion planning, grasp planning, simulation and visi

Open Source Radar — September 22: the self-hosted swap

The Stack

Open Source Radar — September 22: the self-hosted swap

Today's board is one argument told five ways: the paid per-seat product and the self-hosted substitute are now the same product, and one of them runs on hardware you already own. AutoClip (Python, ~8,600 stars, MIT) — Long podcasts, interviews, lectures and stream replays go in; a scored outline, a topic timeline and short clips with generated titles come out. It finds highlights from the transcript rather than watching the video, works from an existing subtitle file or transcribes locally, and

Glasshouse ships an agent-memory benchmark with no headline number

The Stack

Glasshouse ships an agent-memory benchmark with no headline number

A memory-infrastructure vendor has published the test that grades its own industry — and written down, in advance, the ways it could cheat it. Wontopos has put the files for glasshouse v0.1 into the open: a long-term memory benchmark for AI systems whose central design decision is that it will not produce a single score. The corpus is 2,847 questions over a conversation running to 1.97 million tokens, in 10 languages, with 50 uncaptioned photographs, and every axis is reported on its own becaus

Fastcrawl wants to be the web layer for AI agents — and it is priced like a utility

The Stack

Fastcrawl wants to be the web layer for AI agents — and it is priced like a utility

Every AI agent that needs to read the web runs into the same wall, and it is not the model. It is the page. A modern news homepage arrives as 400 to 900 kilobytes of markup wrapped around a few thousand words of actual text — navigation, adverts and script tags around the part worth reading. Handing that to a model wastes tokens and produces worse answers than the same text would have. So the operator either builds a scraping stack — headless browsers, session handling, retries, anti-bot evasion

Cloudflare ships Python Workers GA with native AI bindings

The Stack

Cloudflare ships Python Workers GA with native AI bindings

Cloudflare has spent about two years getting Python onto its edge network. As of Monday it is telling developers to build on it for real. Cloudflare has moved Python Workers to general availability, making Python a first-class language on its developer platform — and, for anyone building AI, giving Python code native bindings to Workers AI, R2, D1, Hyperdrive, Durable Objects, Queues and Workflows instead of forcing conversions into TypeScript at the RPC boundary. That last detail is the one th

Hugging Face ships tokenizers v1 — encoding up to 30x faster

The Stack

Hugging Face ships tokenizers v1 — encoding up to 30x faster

The plumbing of AI got two small releases today — one that stops GPUs from idling on CPU work, one that puts image generation on a laptop-grade footprint. Hugging Face shipped the first major version of its tokenizers library, and the headline number is real: 3 to 30 times faster encoding than v0.23 on a single thread. The tokenizer is the step that turns text into the integer IDs a model reads, and it runs in four stages — normalization, pre-tokenization, the model stage that maps pre-tokens t

M5 Ultra reviews: local AI's ceiling is memory, not compute

The Stack

M5 Ultra reviews: local AI's ceiling is memory, not compute

Apple's local-AI pitch finally met a measurement harness this weekend. The verdict is not about how fast the GPU draws tokens — it is about how fast the machine can read, and how much it can hold while it reads. Federico Viticci spent four days running the top M5 Ultra Mac Studio with 256 GB of memory against an M3 Ultra with 512 GB and his own RTX 5090 desktop, on identical model files and the same runtime. Tom's Hardware, PCMag and Tom's Guide published their own reviews the same day, and the

How to — cut your LLM bill without switching models

The Stack

How to — cut your LLM bill without switching models

Your token spend keeps climbing, and the obvious fix — swap in a cheaper model — is the one you should try last. In a live AI feature, most of the waste isn't the price per token. It's paying full price for tokens the provider has already seen, and generating output tokens nobody asked for. Work the list in this order. The first three moves usually move the number more than a model swap would, and none of them cost you quality. 1. Find out where the tokens actually go. Before changing anythin