Research library
Software engineering
Practical architecture and development decisions for software that people can use, maintain, and adapt.
58 articles · Page 2 of 5
DeepSeek Open-Sourced the Harness That Graded Its Own Benchmarks
DeepSeek Harness (dsh) is MIT-licensed and model-agnostic. It is also the software that produced V4-Pro's Terminal-Bench and DeepSWE scores.
OpenAI's Daybreak Red Ships a Model That Answers 95% of Exploit Requests. Blue Gets 2%.
OpenAI split Daybreak into Blue and Red tiers and gated GPT-5.6-Cyber behind Red. The gap between the tiers says where model refusals actually live.
Meta's Muse Code Cuts Your Token Bill 90%. The Currency Is Your Repository.
Muse Code ships with a contributor tier that cuts token prices over 90% if Meta can train on your code. What that trade is actually worth per developer.
DeepSeek Retrained a 284B Model Until It Beat Its Own 1.6-Trillion Flagship
DeepSeek's official V4-Flash keeps its April architecture and now beats V4-Pro-Preview on all nine agent benchmarks. Post-training did all of it.
Claude Opus 5: Anthropic Undercuts Its Own Flagship
Claude Opus 5 ships at $5/$25, half of Fable 5, with near-Fable scores, looser cyber classifiers and Anthropic's lowest misalignment audit score yet.
Apple Sues OpenAI Over Hardware Trade Secrets — Not Siri, Not Models
Apple Inc. v. Liu is a Defend Trade Secrets Act case about metal finishing, batteries, and suppliers — not Siri, not models. What the docket says.
IBM's Worst Day Ever Is a Warning About What AI Is Doing to Hardware Prices
IBM fell 25.2% on a Q2 miss it partly blamed on clients front-running server and memory price hikes. What the AI buildout did to hardware costs.
Grok Build Was Uploading Whole Git Repos — Secrets Included
A wire analysis showed Grok Build uploading whole Git repos — history and committed secrets — regardless of what the agent read. What to rotate.
Grok 4.5: SpaceXAI's First Coding Model, Pitched as 'Opus-Class'
SpaceXAI's Grok 4.5 is its first coding and agentic model, sold on token efficiency and a $2/$6 price. Musk walked the Opus-class claim back to 4.7.
Anthropic Redeploys Claude Fable 5 After 18 Days Under U.S. Export Controls
Anthropic restored Claude Fable 5 on July 1 after Commerce lifted an 18-day export-control hold triggered by an Amazon jailbreak. What the outage reveals.
Claude Sonnet 5: Anthropic Cuts the Price and Closes the Gap to Opus
Claude Sonnet 5 ships at $2/$10 launch pricing as the most agentic Sonnet yet. The real story: the gap to Opus 4.8 and a new tokenizer.
How AI Finds Vulnerabilities: Fuzzing, Reasoning, and Validation
How AI agents find software vulnerabilities: hybrid fuzzing, static analysis, LLM reasoning, validation, benchmarks, and patch workflows.
