Research library
Insights
Analysis of AI, automation, and software: what is changing, what holds up, and what it means for practical decisions.
155 articles · Page 2 of 13
Anthropic Flipped Claude Code to Auto Mode Because You Were Approving 97% of Prompts Anyway
Claude Code auto mode is now the default on Pro, Max and Team. Its study says humans caught 13.6% of bad commands. What to audit before trusting it.
DeepSeek Open-Sourced the Harness That Graded Its Own Benchmarks
DeepSeek Harness (dsh) is MIT-licensed and model-agnostic. It is also the software that produced V4-Pro's Terminal-Bench and DeepSWE scores.
Qwen3.8-Max's Weights Are Here. The License Isn't Apache and the Checkpoint Isn't the Product.
Alibaba's Qwen3.8-Max weights are live: a text-only, thinking-only checkpoint under a custom license with a $50M commercial gate. Read both before hosting.
The Median Business Spends $12 Per Employee on AI. The Median Worker Saves Under Two Hours.
Ramp's August AI Index and the Census Bureau's worker survey landed a day apart. Both say adoption is broad, spend is cheap, and time saved is modest.
Anthropic Is Watermarking Every Word Claude Writes. It Can't Tell You Who Wrote It.
Anthropic now watermarks Claude's text output worldwide. What the invisible mark actually proves, what strips it, and why detection isn't live yet.
Nvidia's Nemotron Pitch Is a Router That Stops Paying Frontier Prices for Easy Subtasks
Nemotron 3.5 Lightning is a 30B MoE with 3B active, but the open NeMo Switchyard router beside it is Nvidia's real product. How it decides, where it fails.
Grok Bot Gives Each Agent Its Own Computer and Your Logins. Cheapest Seat Is $120 a Month.
SpaceXAI's Grok Bot puts always-on agents inside your apps by signing in as you. What it does, what it costs, and who should wait.
Meta Went Back to Open Weights. The Model Fits in 17 Gigabytes.
Meta open-weighted Muse Glimmer under Apache 2.0: a dense 30B that quantizes to 17GB, runs on a 24GB card, and gives up about 1% of its score.
OpenAI's Daybreak Red Ships a Model That Answers 95% of Exploit Requests. Blue Gets 2%.
OpenAI split Daybreak into Blue and Red tiers and gated GPT-5.6-Cyber behind Red. The gap between the tiers says where model refusals actually live.
Three Labs' Models Reached Outside Their Test Harness in 15 Days. The Sandbox Was the Bug Each Time.
AISI logged 19 unsanctioned agent actions in a cyber eval; Meta's Muse Spark 1.1 hit a third-party service. Four disclosures, one failure: sandbox egress.
Google Restructured DeepMind and Lost Four Researchers on the Same Day
Hassabis becomes chair, Kavukcuoglu takes over, and Jeff Dean leaves with three co-founders for Discovery Loop. What Google's AI reshuffle signals.
Meta's Muse Code Cuts Your Token Bill 90%. The Currency Is Your Repository.
Muse Code ships with a contributor tier that cuts token prices over 90% if Meta can train on your code. What that trade is actually worth per developer.
