Alibaba's CUDA alternative opens to outside developers

Share
Alibaba's CUDA alternative opens to outside developers

Alibaba spent its Apsara conference showing off silicon. The more consequential announcement came the day after, when it opened the software that decides whether anyone can actually use it.

T-Head, Alibaba's chip subsidiary, published a fresh round of open-source releases for T-Head SAIL on September 23 — the CUDA-style stack that sits between PyTorch and its homegrown Zhenwu accelerators — one day after unveiling the Zhenwu V900, which Alibaba CEO Wu Yongming called the most compute-capable AI chip in China at three times the performance of the M890 it succeeds. The new code covers PyTorch adaptation, a source-migration tool called sailify, a Triton-based operator toolchain, and acceleration libraries including DeepGEMM-for-sail and FlashAttention-for-sail. TensorFlow and JAX adapters, an in-house inference engine, and the PCCL and DeepEP communication components are still in progress. SAIL — short for Seed of AI Library — was first open-sourced at WAIC in July, and T-Head says the Zhenwu line now serves more than 650 customers across over 20 industries.

The reason to care is that the software layer, not the transistor count, is what keeps AI teams on Nvidia. T-Head's software-ecosystem director Lu Shenghua framed the customer question plainly: the biggest concern is how much migration costs. So the tools target what teams already have — models written for GPUs, operators tuned over years, hard-won performance knowledge — rather than asking them to start over. Ant Group's inference team has adapted major frontier models to run on the Zhenwu 810E and M890, and is still working through quantization, separating prefill from decode, expert-parallel load balancing and sparse attention. XPeng moved autonomous-driving training from GPU platforms onto Zhenwu cloud clusters, then used SAIL's profiling tools to find its bottlenecks. Xiaohongshu went further, building its own model-migration and operator-optimization agent on SAIL's open code to speed up generative-recommendation deployment — the outcome T-Head says it wants, where customers write their business knowledge back into the toolchain.

There is a second, quieter motive. Custom operators are where a customer's algorithmic edge lives, and few teams want to hand those details to a chip vendor so it can tune their code for them. Opening the stack lets companies write kernels against published architecture information and keep the core logic in-house, while contributing the parts they are willing to share upstream to PyTorch, vLLM and Triton — which also trims the maintenance bill of carrying a private fork. Day-one model support comes with it: by September 2026, T-Head had posted 39 quantized models to ModelScope covering the Qwen, DeepSeek and Kimi families, with more than 348,000 cumulative downloads. That is the number that decides whether a new model is testable on domestic silicon at all.

The strategy is coherent. The constraint has not moved: as we argued this week, Alibaba wants to own the whole stack. Memory says no — no toolchain fixes HBM yields, and the V900 needs memory China cannot yet buy in the volume the plan assumes.

What to watch: whether the PyTorch and vLLM contributions clear upstream review rather than living as a private fork, and whether the V900 ships in volume in the first quarter of 2027 as promised.

Would an open toolchain be enough to move your team off CUDA? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t