Research library

Software engineering

Practical architecture and development decisions for software that people can use, maintain, and adapt.

58 articles · Page 2 of 5

Insight9 min read

DeepSeek Open-Sourced the Harness That Graded Its Own Benchmarks

DeepSeek Harness (dsh) is MIT-licensed and model-agnostic. It is also the software that produced V4-Pro's Terminal-Bench and DeepSWE scores.

Insight7 min read

OpenAI's Daybreak Red Ships a Model That Answers 95% of Exploit Requests. Blue Gets 2%.

OpenAI split Daybreak into Blue and Red tiers and gated GPT-5.6-Cyber behind Red. The gap between the tiers says where model refusals actually live.

Insight9 min read

Meta's Muse Code Cuts Your Token Bill 90%. The Currency Is Your Repository.

Muse Code ships with a contributor tier that cuts token prices over 90% if Meta can train on your code. What that trade is actually worth per developer.

Insight6 min read

DeepSeek Retrained a 284B Model Until It Beat Its Own 1.6-Trillion Flagship

DeepSeek's official V4-Flash keeps its April architecture and now beats V4-Pro-Preview on all nine agent benchmarks. Post-training did all of it.

Insight7 min read

Claude Opus 5: Anthropic Undercuts Its Own Flagship

Claude Opus 5 ships at $5/$25, half of Fable 5, with near-Fable scores, looser cyber classifiers and Anthropic's lowest misalignment audit score yet.

Insight8 min read

Apple Sues OpenAI Over Hardware Trade Secrets — Not Siri, Not Models

Apple Inc. v. Liu is a Defend Trade Secrets Act case about metal finishing, batteries, and suppliers — not Siri, not models. What the docket says.

Insight9 min read

IBM's Worst Day Ever Is a Warning About What AI Is Doing to Hardware Prices

IBM fell 25.2% on a Q2 miss it partly blamed on clients front-running server and memory price hikes. What the AI buildout did to hardware costs.

Insight9 min read

Grok Build Was Uploading Whole Git Repos — Secrets Included

A wire analysis showed Grok Build uploading whole Git repos — history and committed secrets — regardless of what the agent read. What to rotate.

Insight7 min read

Grok 4.5: SpaceXAI's First Coding Model, Pitched as 'Opus-Class'

SpaceXAI's Grok 4.5 is its first coding and agentic model, sold on token efficiency and a $2/$6 price. Musk walked the Opus-class claim back to 4.7.

Insight7 min read

Anthropic Redeploys Claude Fable 5 After 18 Days Under U.S. Export Controls

Anthropic restored Claude Fable 5 on July 1 after Commerce lifted an 18-day export-control hold triggered by an Amazon jailbreak. What the outage reveals.

Insight5 min read

Claude Sonnet 5: Anthropic Cuts the Price and Closes the Gap to Opus

Claude Sonnet 5 ships at $2/$10 launch pricing as the most agentic Sonnet yet. The real story: the gap to Opus 4.8 and a new tokenizer.

Technical guide14 min read

How AI Finds Vulnerabilities: Fuzzing, Reasoning, and Validation

How AI agents find software vulnerabilities: hybrid fuzzing, static analysis, LLM reasoning, validation, benchmarks, and patch workflows.