Huawei open-sources Pangu 2.0's training stack, code and all

Share
Huawei open-sources Pangu 2.0's training stack, code and all

Three things worth your time this hour: the first Chinese lab to give away the code that produces a frontier model, the platforms quietly replacing Hugging Face inside China, and a small OCR model that beat the big generalists on their own benchmark — a month before anyone announced it.

Huawei has open-sourced the code that trains Pangu 2.0, not just the weights — publishing the pretraining, supervised fine-tuning and post-training reinforcement-learning pipelines for the model family. The Sept 28 drop adds two Apache-2.0 repositories alongside the stack Huawei already released: the release notes describe a unified framework covering pretraining, fine-tuning and continued model optimisation, plus a separate reinforcement-learning module that runs on Huawei's Ascend hardware. Both are visible and live at the ascend-tribe organisation Huawei publishes from.

The important detail is what most open releases leave out. Chinese labs routinely publish weights and inference code — DeepSeek, Alibaba's Qwen and Moonshot have all done so, and we covered one such release this week in Xiaomi releases MiMo-V2.6 and open-sources a 1T-class flagship. The pretraining and RL code is the part almost nobody ships, because it is the part that encodes how a model was actually made. Huawei published that, and published it under a permissive licence. This lands directly on the strategy Huawei's rotating chairman laid out at its September developer event — compute is the core bet, and the model layer exists to feed an open hardware ecosystem — and it is the clearest test yet of whether Ascend can attract training work by giving away the recipe rather than the result. We covered the software half of that push when DeepSeek ported its kernels to Huawei's Ascend 950; this is the same bet made by the hardware vendor itself.

Two caveats worth keeping. The repositories are brand new and thin — a single commit each, double-digit stars — so this is a code drop rather than a framework with an external user base; whether anyone reproduces the training run is the only test that counts. And the efficiency figure circulating with the story, a 30% training-efficiency gain when running natively on Ascend, comes from Huawei's own technical report and has not been independently measured. Treat it as a vendor claim until someone outside Shenzhen runs it.


Alibaba's ModelScope and a rival from OSChina are now the two serious answers to Hugging Face inside China — 170,000-odd open models against roughly 20,000 — and neither competes on being the better product. A Rest of World report by Viola Zhou sets out the split: ModelScope, launched in 2022 and run by Alibaba, has grown from about 70,000 models to more than 170,000 in nine months; MoArk, run by the open-source community OSChina, hosts a much smaller library and pitches itself as the one that makes sure mainstream open models run on Chinese chips. The stated motive on both sides is availability, not quality. As OSChina's chief executive Xu Yong told the outlet, the lesson of the internet era was that China developed an independent ecosystem later than it needed to.

The context those decisions are made in matters more than the leaderboard. Hugging Face has been intermittently inaccessible in China since 2023, and Nvidia agreed to acquire it for $12.9 billion this month — a US company buying the default distribution layer for open models. Read against that, ModelScope's growth is not a product-success story so much as an insurance policy. Independent developers quoted in the piece make the harder point: global platforms are still where the work gets discovered, and a domestic mirror has to explain why a Chinese researcher should publish there instead. If you quote a user number for ModelScope, use 25 million — that is what the platform's own reporting says, and the larger figure making the rounds is a translation error.


China Telecom's 1.2B-parameter document parser is genuinely first on the benchmark it claims — which is more than most vendor benchmark claims survive — but it shipped in August, not this week. TeleOCR takes the top spot on OmniDocBench, the document-parsing leaderboard run by the independent OpenDataLab group, with an overall score around 96.9 against roughly 96.5 for the nearest rival, overtaking much larger general vision models including Gemini 3 Pro's 92.9 and GPT-5.2's 86.5 on the same task. Headline comparisons like that flatter a specialist model, and the margin at the top is under a point, but the placement is confirmed on the independent leaderboard rather than only in the vendor's own table — and the weights are public under a permissive licence, so anyone can check. The date is the part to correct: the paper and the model went up in mid-August, the leaderboard entry landed in September, and the wire release this week is a paid announcement. The model is real and it is open. It is simply not news today, and framing a repackaged release as a launch is the kind of thing readers eventually notice.

What to watch: whether anyone outside Huawei reproduces the Pangu training run from the released code. A recipe nobody cooks is a press release; a recipe that works is an ecosystem.

Huawei gave away the training code but not the hardware to run it — is open-sourcing a pipeline still open if it only really runs on one vendor's chips? Tell us in the comments.

Sources: OmniDocBench leaderboard (OpenDataLab)

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t