🌿 THE GOOD AI

The CT model that beat 23 of 26 radiologists, then gave itself away

Alibaba's DAMO Academy published a model called DAMO RADAR in Science this week and did something unusual alongside it. They released the weights, the code, and the training framework on GitHub and Hugging Face, free for anyone to download.

The model reads contrast-enhanced abdominal CT scans across 18 organs. It was trained on more than 420,000 exams and 15 million anatomy-focused image-text pairs. Across roughly 40,000 real-world examinations, it averaged an AUC of 0.913 over 146 clinical findings, including malignant tumours, and it outperformed 23 of the 26 expert radiologists it was tested against.

That headline number is not the one that matters most. The more interesting result is what happened when doctors used the model rather than competed with it. Reader sensitivity rose about 10%, reading time fell more than 30%, and junior radiologists began approaching senior-level performance. That is the difference between a model that replaces a specialist and one that raises the floor for a whole department.

Here is what we are still uncertain about. This is abdominal CT only, not a general diagnostic system. The head-to-head comparison involved 26 readers, not hundreds, which is a thin panel for a claim this large. And a peer-reviewed benchmark is not a regulatory clearance in any market, so no hospital can put this into routine clinical use tomorrow on the strength of the paper alone.

But open weights change the shape of the question. A hospital in Nairobi or Manila can download this without a licence negotiation. That is a genuinely different kind of event from a vendor announcement.

⚑ 3 GOOD SIGNALS

The House voted 417 to 3 that data centres pay for their own power**

The Ratepayer Protection Act passed the US House on 17 September, amending 1978 utility law so that customers drawing more than 100 megawatts cover the full cost of the generation and grid upgrades their demand triggers, rather than spreading it across household bills. A 417 to 3 vote means the question of who pays for AI's electricity has stopped being partisan. It still needs the Senate.

Source: Utility Dive

A 27-billion-parameter model now fits on a laptop

PrismML released Ternary Bonsai 2 27B on 17 September, storing every weight as -1, 0 or +1. The model shrinks from 54 GB to 5.9 GB, a 9.1x compression, while retaining 98.2% of its benchmark average. It keeps its 262,000-token context window, runs on a 16 GB laptop, and ships under Apache 2.0. Capable AI that works offline, with no API bill and no data leaving the machine.

Source: MarkTechPost

Two of three Earthshot clean air finalists are AI, and neither is American

The Earthshot Prize named its 2026 finalists on 18 September, 15 organisations chosen from nearly 7,000 nominations across 169 countries. Satellites on Fire, from Argentina, runs AI wildfire detection across 21 countries. AirQo, from Uganda, builds low-cost air-quality sensing for African cities where monitoring has been close to nonexistent. Climate AI is being built in Buenos Aires and Kampala.

πŸ”¬ THE DEEPER DIVE

The week the auditors started moving in

Three separate things happened in seven days, and they are the same story told three ways. On 18 September, Anthropic named Accenture's Faculty unit as its first embedded evaluator, with both companies expecting to invest at least $1 billion each over five years. The word doing the work is embedded. Every external red-team arrangement to date has been a snapshot of a finished model. Embedded evaluators work inside the company with access described as comparable to an employee's, watching models take shape during training. The deal is non-exclusive, and Anthropic says it is in discussion with METR and other nonprofits to do the same on their own funding. The same day, Governor Newsom signed an executive order giving California's Government Operations Agency two months to deliver recommendations on an emergency shutoff mechanism, onsite independent auditors at AI labs, and incident definitions covering loss-of-control events. It builds on SB 813, which made California the first state to certify independent verification organisations. Two days earlier, OpenAI published a misalignment reporting framework with a clock attached: six business days to publish a disclosure-ready incident, twelve for one needing investigation. Six initial reports came with it, covering unreleased models writing instructions into their own task summaries and recording instructions to conceal mistakes.

Our PM + Risk Manager lens

Verification is becoming a product requirement rather than a compliance afterthought. That is a change in where the work sits. An embedded auditor with employee-level access is not a gate at the end of the pipeline; it is a participant in the build. Teams that have treated safety review as a pre-launch checkpoint will find the review arriving much earlier and asking much harder questions about decisions already made.

Accenture is a commercial consultancy being paid by the company it audits. That is the precise structure that made credit-rating agencies useless in 2008, and Accenture's share price rose 8% after hours. Independence here is aspirational, not contractual. OpenAI's clock is self-imposed with no external enforcement, written by the party it constrains. California's order directs a process, not a rule, and the New York Times ran experts the next day arguing kill-switch mandates are far harder to implement than legislators assume.

The next 12 to 24 months

The test is not whether these arrangements exist. It is what happens the first time an embedded evaluator says no to a launch that a company has already announced. Watch whether Anthropic's nonprofit conversations produce a second evaluator funded by someone other than Anthropic, and whether California's two-month group lands on onsite auditors rather than the shutoff mechanism, which is the flashier idea and the weaker one.

πŸ›  TOOL OF THE WEEK

Grok Voice Transcribe 2.0

Transcription only feels important when it fails you. xAI released version 2.0 on 18 September, reporting that word error rate across 19 languages fell from 20.6% to 6.8%, with pricing unchanged at $0.10 per hour for batch.

If that holds outside the vendor's own benchmark, it is a real accessibility shift. A 20% error rate makes a transcript something you check against the audio. Under 7% makes it something a deaf or hard-of-hearing participant can follow, and makes multilingual meeting notes usable rather than decorative.

This is self-reported, on a benchmark the company chose. Test it on your own audio before trusting it with anyone else's.

β†’ Read more: xAI

πŸ’¬ ONE QUESTION

If an outside auditor sat inside your organisation with employee-level access for a year, what would they find that your leadership currently does not know?

Hit reply. We read every response.