🌿 THE GOOD AI

An AI tutor just passed education's hardest test

Google DeepMind, education nonprofit Fab AI, and Sierra Leone's Ministry of Education ran a randomized controlled trial, the same evidence standard used for new medicines, on Gemini's Guided Learning mode. For eight weeks, 1,763 junior secondary students across 12 schools in Port Loko District used the tutor in math class. Students with access gained 0.258 standard deviations over the control group, roughly 1.2 to 1.7 years of typical learning progress. In classrooms where teachers built Gemini into about half their lessons, hitting a 12-hour usage target, gains ran higher still, roughly 1.8 to 2.5 years' worth.

Education technology has a long habit of glowing pilots that evaporate under rigorous testing, and engagement is usually the first thing to collapse. Voluntary edtech tools typically see about 5 percent of students meet usage targets. Here, 69 percent did. And it happened in a school system with a severe teacher shortage, which is exactly where an always-available tutor matters most. This is some of the first gold-standard evidence that a general-purpose AI tutor can move real learning outcomes at classroom scale.

This is one subject, one district, eight weeks, and Google evaluating Google. Results were published in June, gains were measured at the trial's end, and nobody yet knows whether they persist or whether independent teams can replicate them elsewhere. The teachers, not the tool alone, look like the active ingredient, and that is harder to scale than software.

⚑ 3 GOOD SIGNALS

Indian doctors report 10 more patients a week, and they are double-checking the AI

The Philips Future Health Index India report found 71 percent of Indian clinicians say AI increased their capacity to see patients, a median of 10 extra per week, while 86 percent insist every AI output needs human review and 78 percent have caught and corrected AI errors. Capacity gains and skepticism, working together.

Source: The Week

California moved 24 AI bills in a single day

At the August 13 suspense votes, appropriations committees advanced 24 of 29 live AI bills to floor votes, covering chatbot safety for children, workplace surveillance, and student privacy. Two have reached Governor Newsom's desk. With 85 AI laws already enacted across 27 states this year, the American AI rulebook is being written now, state by state.

Three North Carolina newsrooms wrote the playbook for AI without losing trust

Public and nonprofit newsrooms in North Carolina are using AI for quizzes, captions, audio cleanup, archive digitization, and transcription, while publishing disclosure policies and keeping humans on every editorial call. Adopt it where it saves labor, disclose it, decide nothing by machine. Any mission-driven organization could copy this today.

Source: Current

πŸ”¬ THE DEEPER DIVE

The manipulation study nobody wanted and everybody needed

In a study titled "Evaluating Language Models for Harmful Manipulation," Google DeepMind paid 10,101 real people across the US, UK, and India, then set a frontier model to advocate positions in public policy, personal finance, and health. The question: could it shift not just stated beliefs but actual spending decisions, with real money on the table? The answer was yes. It is the largest empirical study of AI manipulation to date, DeepMind ran it on its own model, and it published a measurement method any lab or regulator can now reuse. The finding that matters most is a negative one. How often a model reaches for recognizable manipulative tactics does not predict how successfully it changes what people do. That quietly breaks the industry's comfortable assumption that you can police manipulation by watching for it, flagging emotional appeals and false urgency and calling the job done. Tactics and outcomes are decoupled. A model can look clean and still move you.

Our PM + Risk Manager lens

Measurement is the product decision here. Persuasive capability is not a feature anyone asked for, but it ships inside every conversational product whether you ordered it or not. The mature move is what DeepMind did: quantify the unwanted capability before someone else quantifies it for you, and publish the unflattering number. Teams building on frontier models should treat manipulation evals the way they treat load testing, a gate you clear before launch rather than a question you answer after an incident. With a reusable protocol now public, "we had no way to test that" stops being an available answer.

The decoupling result should reorganize control frameworks. Most manipulation controls today are detection-based: monitor outputs for manipulative language, block what you catch. This study says that control does not map to the risk. The exposure lives in outcomes, changed beliefs, and changed spending, not in vocabulary. That points toward outcome-based testing with real participants, third-party audits, and red lines defined by effect size rather than word lists. One discipline note: resist any celebratory framing. This is a capability nobody wants, and the win is only that it is now measurable.

The next 12–24 months

Watch three things. Whether other labs run this protocol on their own models and publish, which would turn manipulation testing from a paper into a norm. Whether regulators, with the EU AI Act now in force and 27 US states legislating, fold outcome-based manipulation evals into their requirements. And whether anyone builds the equivalent for agentic systems, which will not just argue a position but act on one. The uncomfortable study is usually the one that ages best.

Sources: arXiv, IBTimes UK

πŸ›  TOOL OF THE WEEK

WeatherNext Cyclones

It is peak Atlantic hurricane season, which makes this the week to know about WeatherNext Cyclones. Google DeepMind's storm model, published in Nature this month, delivers three-day cyclone track and intensity forecasts with the accuracy older systems managed at two days. That is an extra day of warning for anyone in the path. The US National Hurricane Center tested it through the 2025 season, including Hurricane Melissa's rapid intensification. DeepMind open-sourced the code and model weights under Apache 2.0, and a mini version runs in a free Colab notebook. An evacuation-grade forecast, given away.

β†’ Read more: Google DeepMind

πŸ’¬ ONE QUESTION

The Sierra Leone trial had a quiet twist: the biggest gains came in classrooms where teachers folded the AI into their lessons, not where students used it alone.

Where have you seen a tool succeed only because of the person who decided how to use it?

Hit reply. We read every response.