🤝 The study group that beat the answer key
Two models with no answer key grade each other's work and out-learn the one trained on real labels. Plus a safety report card and a DNA test for weights.
The desk behind AI Reach-out Frontier. We read the whole paper so you don't have to pretend you did.
Two models with no answer key grade each other's work and out-learn the one trained on real labels. Plus a safety report card and a DNA test for weights.
Providers hand you an encrypted copy of a model's reasoning. A paper feeds it to a weaker sibling and reads it aloud. Plus the week's research news.
One engineer let an AI rewrite 189 files of a 717k-line app with zero code review — the trick was auditing the spec instead. Plus the week's research news.
A 150M-parameter model scores 29.5% on ARC-AGI for $0.0007 a task — by thinking in silence. Latent reasoning in plain words, plus the week's research news.
AudioRubrics, plainly: teach an audio model with a rubric that evolves mid-training. Plus Qwen3.8-Max, DeepSeek V4-Flash, and Anthropic on open weights.
A new paper builds memory directly into the model instead of bolting notes on the side. We explain Metis in plain words, plus the week's research news.
News
The research behind the headlines, in plain words. One paper explained and the news that matters most.