Research library

Artificial intelligence

Evaluate where AI is useful, understand the systems behind it, and make informed choices about models, data, and deployment.

158 articles · Page 2 of 14

Insight10 min read

DeepSeek Shipped V4-Pro, Then Raised Its Prices

DeepSeek V4-Pro went GA Aug 13, then API prices rose 1.5x to 12x on Aug 16 with peak/off-peak tiers. Old vs new rate table and the cost-planning fallout.

Insight9 min read

Anthropic Flipped Claude Code to Auto Mode Because You Were Approving 97% of Prompts Anyway

Claude Code auto mode is now the default on Pro, Max and Team. Its study says humans caught 13.6% of bad commands. What to audit before trusting it.

Insight9 min read

DeepSeek Open-Sourced the Harness That Graded Its Own Benchmarks

DeepSeek Harness (dsh) is MIT-licensed and model-agnostic. It is also the software that produced V4-Pro's Terminal-Bench and DeepSWE scores.

Technical guide13 min read

One Shared Key Let Anyone Decode What Claude, GPT and Gemini Were Hiding in Their Reasoning

One shared key sealed the reasoning blocks Anthropic, OpenAI and Google return to clients, letting a weaker model decode a stronger sibling's thoughts.

Insight10 min read

Qwen3.8-Max's Weights Are Here. The License Isn't Apache and the Checkpoint Isn't the Product.

Alibaba's Qwen3.8-Max weights are live: a text-only, thinking-only checkpoint under a custom license with a $50M commercial gate. Read both before hosting.

Insight10 min read

The Median Business Spends $12 Per Employee on AI. The Median Worker Saves Under Two Hours.

Ramp's August AI Index and the Census Bureau's worker survey landed a day apart. Both say adoption is broad, spend is cheap, and time saved is modest.

Insight9 min read

Anthropic Is Watermarking Every Word Claude Writes. It Can't Tell You Who Wrote It.

Anthropic now watermarks Claude's text output worldwide. What the invisible mark actually proves, what strips it, and why detection isn't live yet.

Insight10 min read

Nvidia's Nemotron Pitch Is a Router That Stops Paying Frontier Prices for Easy Subtasks

Nemotron 3.5 Lightning is a 30B MoE with 3B active, but the open NeMo Switchyard router beside it is Nvidia's real product. How it decides, where it fails.

Insight8 min read

Grok Bot Gives Each Agent Its Own Computer and Your Logins. Cheapest Seat Is $120 a Month.

SpaceXAI's Grok Bot puts always-on agents inside your apps by signing in as you. What it does, what it costs, and who should wait.

Insight9 min read

Meta Went Back to Open Weights. The Model Fits in 17 Gigabytes.

Meta open-weighted Muse Glimmer under Apache 2.0: a dense 30B that quantizes to 17GB, runs on a 24GB card, and gives up about 1% of its score.

Insight7 min read

OpenAI's Daybreak Red Ships a Model That Answers 95% of Exploit Requests. Blue Gets 2%.

OpenAI split Daybreak into Blue and Red tiers and gated GPT-5.6-Cyber behind Red. The gap between the tiers says where model refusals actually live.

Insight10 min read

Three Labs' Models Reached Outside Their Test Harness in 15 Days. The Sandbox Was the Bug Each Time.

AISI logged 19 unsanctioned agent actions in a cyber eval; Meta's Muse Spark 1.1 hit a third-party service. Four disclosures, one failure: sandbox egress.