πΏ THE GOOD AI
AI found 44 sperm cells in a sample two technicians had searched for two days
Columbia University Fertility Center calls it STAR, for Sperm Tracking and Recovery. The system pushes a semen sample through a microfluidic chip thinner than a human hair while a high-speed camera captures about 300 images per second, more than 8 million images per sample. AI scans the flow, flags individual sperm, and diverts each one into a capture valve. In the case that made headlines this week, a sample that two embryologists had searched for two days without success yielded 44 viable cells in about an hour.
Across 175 cases of suspected azoospermia, a condition affecting roughly 10 percent of infertile men, STAR has recovered viable sperm 26 percent of the time, and the first pregnancy from the method was reported last fall. Male infertility contributes to up to 40 percent of infertility cases worldwide, and for men whose samples appear empty, the options have been surgery or nothing. This is AI doing something humans physically cannot: searching millions of cellular fragments without fatigue and without ever deciding the search is hopeless. The results are measured in families rather than benchmarks.
STAR exists at one hospital, on three machines, and it cannot help men who produce no sperm at all. This week's reporting is the first look at the full 175-case record, and a 26 percent recovery rate is a real opening, not a guarantee. Fertility medicine has a long history of promising technology meeting complicated biology.
Source: The New York Times
β‘ 3 GOOD SIGNALS
Claude's output now carries an invisible signature, worldwide
Anthropic published how it watermarks everything Claude generates: the C2PA provenance standard for files, plus a hidden mark in text that survives copy and paste. The EU AI Act requires this in Europe only; Anthropic applied it globally, with detection tools promised. A mark means Claude processed the content, not that no human was involved.
Source: Anthropic
The workers AI supposedly replaces are its heaviest users
OpenAI's analysis of roughly 17 million ChatGPT Enterprise messages found early-career employees send 8 to 9 more messages per week than executives, and agentic work is now 64 percent of enterprise output tokens. OpenAI also published the awkward part: no meaningful correlation between usage and revenue per employee. The bottom rung of the ladder is climbing hardest.
A monkey detector in Nepal knows whether the macaque is watching or already eating
At Madan Bhandari University, researcher Progress Jung Thapa trained a lightweight edge-computing camera on more than 4,000 images to spot crop-raiding macaques and classify their behaviour, moving, feeding, or vigilant, reaching 88 percent accuracy across 28 field trials before pinging farmers' phones. Villagers currently guard fields in two-hour shifts. His line: "A machine can detect. But a human decides."
Source: Mongabay
π¬ THE DEEPER DIVE
Sixty agents, 650 dead ends, and a problem posed in 1859
On August 10, Anthropic disclosed that an unreleased research version of Claude raised the proven lower bound for the share of Riemann zeta zeros on the critical line from 41.6 percent to 67.2 percent, the largest single advance in the problem's history. In the 37 years before this, human mathematicians had moved that bound by less than one percentage point. The system coordinated roughly 60 subagents, burned through about 31 million output tokens, and discarded some 650 failed ideas before finding the move that worked: combining two published papers no mathematician had thought to connect. The proof was formally verified in Lean 4 and reviewed by two external mathematicians, and the code and process logs are public. Anthropic is plain that the approach is not expected to prove the full Riemann Hypothesis, which remains open.
Our PM + Risk lens
The product lesson is the architecture, not the theorem. Claude did not out-think the field; it out-searched it, and it could afford 650 failures because checking was automatic. Lean 4 gave the system a machine-checkable definition of done, which turns wild generative breadth from a liability into a strategy. That pattern ports to any domain with a hard verifier: test suites, formal specs, chip design rules. The other decision worth copying is publishing the process logs. A result is a demo. A documented, reproducible pipeline is a method, and methods are what other teams adopt.
Notice what nobody is asked to do here: trust the model. The claim rests on a mechanical proof checker, two independent human reviewers, and public logs, layered controls that hold even if one fails. Add the explicit scoping that this will not crack the full hypothesis, and you have what claim discipline looks like. The risk sits downstream of the headline. Mathematics has Lean; most fields that will borrow the "AI research partner" story have no verifier at all. Expect confidence to travel faster than verification infrastructure, and treat machine-scale results without a checking layer as press releases until proven otherwise.
The next 12β24 months
This is the third significant AI-assisted mathematics result of the summer, which is starting to look like a pattern rather than a stunt. Watch whether the swarm-plus-verifier recipe lands where machine-checkable correctness already exists, in software verification, chip design, and parts of drug discovery. Watch whether journals and prize committees build review processes for proofs too large for any one human to read. And watch whether public logs become the norm. If they do, the interesting question stops being whether AI can do research and becomes who gets to check it.
Sources: Forbes,
π TOOL OF THE WEEK
SPARROW
Conservation's bottleneck has rarely been analysis. It is collection: getting data from places humans struggle to reach. SPARROW, from Microsoft's AI for Good Lab, is a solar-powered listening post that gathers images and audio in remote forests, runs its AI on the device itself, and relays findings home without anyone making the trip. Deployed units have run for more than a year without interruption, across five continents. As lab director Juan Lavista Ferres puts it, "it takes a huge amount of effort to collect data." This is that effort, automated, so conservation models finally have something to learn from.
β Read more: Mongabay
π¬ ONE QUESTION
STAR did not out-think the embryologists. It outlasted them, searching millions of images without tiring and without losing hope. Where in your life or work is the real limit not intelligence, but stamina?
Hit reply. We read every response.
