The Stack

Software, APIs, and things you can use

Open Source Radar — September 28: the labs open their internal tools

The Stack

Open Source Radar — September 28: the labs open their internal tools

Today's signal is vendors publishing the tools they built for their own engineers — Alibaba's review bot, Tencent's knowledge platform, Cloudflare's audit skill, Anthropic's role plugins — while a one-person project bets that the right answer to Claude Code versus Codex is both. Open Code Review (Go, Apache-2.0, ~42,000 stars) — Alibaba ran this internally as its official AI code review assistant for two years before incubating it in the open, and the case for that decision is measurable. It re

Quantum work goes local — a desktop box, an agent as the interface

The Stack

Quantum work goes local — a desktop box, an agent as the interface

Today's most interesting quantum launch is not about qubits — it is about who, or what, sits between a scientist and the machine. Also: the anonymous model everyone is benchmarking hit the top of the usage charts, and China's "quantum-enhanced LLM" pitch is turning into a queue. Unitary Quantum, a Shanghai Jiao Tong University spinout, launched UnitarySpark on September 20 — a desktop workstation that puts quantum simulation, GPU acceleration, an AI agent and a link to real quantum processors i

Open Source Radar — September 27: agents get a workplace

The Stack

Open Source Radar — September 27: agents get a workplace

Today's trending list is mostly about where agents work rather than what they think with: an org chart, a fleet console, a team room, and two new ways for them to reach phones and pull requests. Paperclip (TypeScript, ~88,200 stars, MIT) — A Node server with a React dashboard that treats a group of agents as a company. You declare a goal, staff it with whatever you already run — Claude Code, Codex, Cursor, OpenClaw, or anything that can hold an HTTP endpoint — set budgets, and watch the work fr

AI 101 — What is prompt caching?

The Stack

AI 101 — What is prompt caching?

Prompt caching is a way to make an AI model skip re-reading text it has already read. The provider saves its internal working state for the opening section of your prompt, and the next request that starts with the exact same text picks up from there — cheaper and faster. It sounds like a developer-housekeeping detail. It is now one of the largest line items in how AI gets priced, and one of the most common ways a working product quietly runs up a bill. Why it matters right now On September 2

llama.cpp's CPU prompt processing gets 3–7x faster on k-quants

The Stack

llama.cpp's CPU prompt processing gets 3–7x faster on k-quants

Three local-AI items for this hour: a kernel change that makes the wait before the first token 3 to 7 times shorter on a plain CPU, two small Chinese models that answer questions instead of writing text, and a 2017 laptop finishing real agent work. Local models just got dramatically faster to query on a processor most people already own — a kernel merged into llama.cpp today makes prompt processing 3 to 7 times quicker for the quantised formats the local community actually runs. The change rewo

Boom loses its launch customer as Crusoe drops $1.25B turbine order

The Stack

Boom loses its launch customer as Crusoe drops $1.25B turbine order

AI's power buildout just got its first visible cancellation, and a 4B model out of the open-source world showed how cheap an agent's decision layer can be. Crusoe has walked away from a $1.25 billion order for Boom Supersonic's Superpower turbines, ending the launch partnership the two Denver companies announced last December. Boom chief executive Blake Scholl broke the news Friday evening on X, writing that "turbines are no longer part of Crusoe's near term primary power mix at Abilene/etc., s

Alibaba's CUDA alternative opens to outside developers

The Stack

Alibaba's CUDA alternative opens to outside developers

Alibaba spent its Apsara conference showing off silicon. The more consequential announcement came the day after, when it opened the software that decides whether anyone can actually use it. T-Head, Alibaba's chip subsidiary, published a fresh round of open-source releases for T-Head SAIL on September 23 — the CUDA-style stack that sits between PyTorch and its homegrown Zhenwu accelerators — one day after unveiling the Zhenwu V900, which Alibaba CEO Wu Yongming called the most compute-capable AI

Open Source Radar — September 25: memory, method, smaller models

The Stack

Open Source Radar — September 25: memory, method, smaller models

Today's trending list splits three ways: what an agent remembers, how it works, and whether any of it fits on hardware you already own. Hindsight (Python, ~28,400 stars, MIT) — Vectorize's memory layer is one of the most-trended repos on GitHub today, and the distinction it draws is the useful part: most agent memory is conversation history, while this is built so an agent's beliefs change as evidence accumulates. Facts go in through one call, get consolidated in the background into observation

Copilot made Okta's engineers faster but not more productive

The Stack

Copilot made Okta's engineers faster but not more productive

A year-long field study inside Okta finds the individual gains are real and the organisational ones are not there. Plus: the Jev team ships a code reviewer built on the assumption that nobody reads a 230-file diff. A peer-reviewed field study inside Okta found that GitHub Copilot cut engineers' working hours and sharply lifted their motivation and perceived skill — while producing no statistically significant change in the code the company shipped. The paper, published in Communications of the

AI 101 — What is federated learning?

The Stack

AI 101 — What is federated learning?

Federated learning is a way to train one shared AI model across thousands of devices or organisations without ever collecting their data in one place — each participant trains a copy on data it already holds, and only a summary of what that training learned travels back. The name comes from the word "federation": a group of independent members cooperating under one agreed protocol while each keeps its own affairs. The data stays where it lives. The learning moves. Why it matters right now Tw

Fastino's 340M decision model beats its own 1B on routing

The Stack

Fastino's 340M decision model beats its own 1B on routing

Most agent stacks answer a yes/no question by having a large model write a paragraph. Fastino's answer is to not ask a large model at all — and its new release is small enough to run on the CPU you already have. Fastino shipped GLiNER2.5-Decide, a 340M-parameter classifier that takes a bounded question — intent, routing, sentiment, severity, spam — and returns a label in a single forward pass, with no prompt template and no generated tokens. It is Apache 2.0, built on a DeBERTa-v3-large encoder

How to — keep your AI feature alive when a model is retired

The Stack

How to — keep your AI feature alive when a model is retired

Every model you call has a shut-down date, and most teams only learn it from a failed request in production. The dates are published months ahead. The breakage happens anyway. Treat retirement as a scheduled maintenance window rather than an emergency: know what you depend on, know your provider's clock, and have the replacement tested before the deadline arrives. 1. Inventory every model identifier you depend on — not just the one in your config. The model name is usually in more places than