🌿 THE GOOD AI

MIT built an AI that can picture the storm nobody has ever recorded

Every tool we currently use to estimate catastrophe risk has the same flaw. It learns from historical examples of the catastrophe, which is a problem when the catastrophe is, by definition, rare. If a city has never seen a given storm, the model has nothing to learn from, and the engineer sizing the drainage has nothing to consult.

MIT researchers published a method in Nature Communications on August 20 that gets around this. Called Extreme Event Aware, it trains on ordinary weather rather than on a catalogue of disasters, learning the underlying shape and physics of rainfall instead of memorising which storms happened to occur. From that, it generates thousands of plausible versions of an event with no entry in the record at all. New York City's heaviest recorded rainfall is roughly 200 millimetres. A planner who wants to design for 300 millimetres can now see what that would look like, block by block.

Flood defences, insurance pricing and municipal capital budgets are all being set right now against a historical record that no longer describes the climate anyone actually lives in. This is the first credible way to plan against the storm that has not happened yet. One of the researchers frames the whole problem in a single question: "What will be the Katrina that happens every 100 years?"

These are generated scenarios, not forecasts. A model trained on ordinary weather is making a physically informed extrapolation, and nobody can validate it against an event that has never occurred. The honest test will be whether planners and insurers treat the output as a planning input or as a prediction.

⚑ 3 GOOD SIGNALS

⚑ An AI stopped copying evolution's homework and got better at designing proteins

Protein design models are usually scored on whether they reproduce the sequence evolution happened to pick. MIT biologist Amy Keating argues that is the wrong target. Her lab's PottsMPNN builds in the physics of protein stability instead, and as it leaned less on natural sequences, its predictions improved. Designed proteins are the front end of the drug pipeline.

Source: MIT News

🀝 AI found six ways to print a NASA rocket alloy on ordinary machines, in 40 tries

GRCop-42, the copper alloy NASA developed for rocket combustion chambers, is notoriously hard to 3D print. Washington State University researchers used an adaptive experimental design model to search over 100 million printer configurations. After 40 physical experiments, it surfaced six that worked, including one at a record-low 500 watts, putting the alloy within reach of far cheaper machines.

Source: WSU Insider

🌱 Researchers hid orders in an email in invisible ink, hijacked the AI summary ten times out of ten, and published the fix

Forcepoint X-Labs planted HTML at zero font size in white text inside an email. The reader saw 537 clean characters. The summarising model received 1,009, including hidden instructions, and every one of ten runs produced a manipulated summary. Forcepoint published the mitigation alongside the finding: extract only user-visible content, and treat email as untrusted data.

Source: CSO Online

πŸ”¬ THE DEEPER DIVE

Two years of data on an AI tutor: the students who had it barely used it

Last week we led with a randomised controlled trial in Sierra Leone where 69 percent of students hit their AI tutoring usage target, against a typical edtech figure of about 5 percent. This week brings the other half of the picture, and it is the more common one.

A two-year school study of Khanmigo, Khan Academy's AI tutor, tracked what students actually did with it. Students on the platform did make faster maths gains than the comparison group, so the tool works when it is used. But the median student messaged the tutor on roughly a third of available days. Researchers described the result as near-universal access paired with thin engagement.

Access is not adoption

The gap between those two studies is the whole story. Sierra Leone's trial built the tutor into scheduled class time with teachers driving usage. The Khanmigo deployment made it available. Availability, on this evidence, is not a strategy.

Our PM + Risk Manager lens

This is a product failure, not a technology failure. The model was good enough to move maths scores. What was missing was the engagement design: the trigger, the habit loop, the reason to open it on a Tuesday. Most AI education products are still shipped as a capability and measured on seats provisioned. The number that matters is median days engaged, and almost nobody is reporting it. If your success metric is licences deployed, you will ship exactly this result.

Districts have signed multi-year contracts against an assumed uptake rate nobody validated. That is a procurement exposure, and it is quietly large. When the efficacy review comes, and the usage data looks like this, the tool takes the blame for what was really an implementation gap, and the budget gets cut for a product that was working. The mitigation is unglamorous: contract on measured engagement thresholds, not on headcount, and instrument usage from day one.

The next 12–24 months

Expect engagement, not accuracy, to become the contested metric in education AI. The vendors who report median days engaged voluntarily will look conservative for about a year and credible after that. We would also expect a wave of district-level results that look like this one, followed by a correction in how these tools get bought. The honest read is that AI tutoring works and deploying it does not.

Source: Chalkbeat

πŸ›  TOOL OF THE WEEK

Claude's memory, with the delete button included

Anthropic merged the memory behind Claude chat and Claude Cowork this week, so context carries between them. The part worth noticing is the controls. You can open what has been stored and edit or delete any of it. Claude files topics as a conversation happens rather than summarising you afterwards. Sensitive categories including health, race, religion, politics and gender identity stay off behind a toggle, and ID numbers and immigration status are never stored at all. It is on by default across Free, Pro and Max. Memory is the feature most likely to quietly build a profile of someone, so shipping inspection and deletion at launch rather than after the first scandal is the bar worth holding everyone to.

β†’ Read more: TechCrunch

πŸ’¬ ONE QUESTION

If an AI tool at your work is available to everyone and used properly by almost nobody, whose problem is that: the tool's, the training's, or the job's?

Hit reply. We read every response.