Research library
Artificial intelligence
Evaluate where AI is useful, understand the systems behind it, and make informed choices about models, data, and deployment.
158 articles · Page 7 of 14
Claude Opus 4.8: Anthropic Bets on Honesty and Subagent Orchestration
Opus 4.8 is 4x less likely to let its own code flaws slide, adds dynamic workflows orchestrating up to 1,000 subagents, and cuts fast mode pricing.
Project Glasswing's First Month: 10,000 Vulnerabilities Found, and Why That Was the Easy Part
Anthropic's Project Glasswing found 10,000+ high-severity flaws in a month via Claude Mythos Preview. Finding bugs got cheap; fixing them didn't.
OpenAI's Codex Update: Goal Mode Graduates and Codex Reaches Off the Terminal
Codex's May 21 update takes Goal Mode out of beta, adds macOS Appshots and remote desktop control, and opens a plugin marketplace to Business users.
Google I/O 2026: Gemini Spark, the AI Ultra Repricing, and the Pro Model That Slipped
Google I/O 2026 shipped Gemini 3.5 Flash, the Spark agent, Android XR glasses, and a $100 AI Ultra tier. The Pro model slipped to June.
TPU v8 vs Blackwell: How AI Silicon Is Splitting Into Training and Inference Chips
Training and inference need different silicon. TPU v8t/v8i architecture, comparison to Blackwell's unified design, and per-token cost implications.
OpenAI Ships Codex in the ChatGPT Mobile App — The Phone Becomes an Agent Remote Control
Codex now runs in the ChatGPT mobile app as a remote control for macOS agents — the phone supervises while code and credentials stay on desktop.
Amazon Replaces Rufus With Alexa for Shopping — Agentic Commerce Enters the Default Retail UX
Alexa for Shopping merges Rufus and Alexa+ into one agent inside the Amazon search bar — a shift from chatbots to agentic commerce at scale.
Anthropic Launches Claude for Small Business — Packaged Workflows for QuickBooks, HubSpot, PayPal
Anthropic's Claude for Small Business ships 15 packaged agentic workflows wired into QuickBooks, HubSpot, PayPal, and Canva. What SMB owners need to know.
Compressed Sparse Attention: How DeepSeek V4 Reached 1M Context at 27% of the FLOPs
DeepSeek V4 hits 1M context at 27% of V3.2's per-token compute. How Compressed Sparse Attention and Heavily Compressed Attention combine to do it.
KV Cache: The Hidden Memory Wall in LLM Inference
The KV cache memory wall in LLM inference: the math behind long context costs and architectural solutions (GQA, MQA, MLA, paged attention).
Microsoft Copilot Studio Goes Multi-Agent — GPT-5.5 Reasoning, Cross-Vendor Governance, Work IQ
Copilot Studio May 2026: chatbot builder to agent governance. Cross-vendor policies, agent orchestration, MCP support, GPT-5.5 reasoning access.
EU AI Omnibus: High-Risk AI Rules Pushed to December 2027
EU delays high-risk AI enforcement 18 months to Dec 2027, accelerates synthetic-content transparency to Dec 2026. Regulatory sandboxes now August 2027.
