Browse guides

AI News articles

34
Sakana AI

Fugu-Ultra v1.1 Released: Sakana AI's Orchestration Model Gets a Frontier Refresh

Sakana AI shipped Fugu-Ultra v1.1 on July 24, 2026, refreshing its multi-agent orchestration model with newer frontier workers. Sakana claims up to 7.9 benchmark points over v1.0, with the biggest gains on ProgramBench and Terminal Bench 2.1, and keeps pricing identical to v1.0. Here is what changed, what is still vendor-reported, and how to evaluate the upgrade.

Claude AI

Claude Opus 5: Benchmarks, Pricing, and Full Guide (July 2026)

Claude Opus 5 launched July 24, 2026. It delivers near-Fable 5 intelligence at Opus pricing ($5/$25 per MTok), with a 1M-token context window, 128K max output, and state-of-the-art results on Frontier-Bench, ARC-AGI 3, and OSWorld 2.0. This guide covers benchmarks, pricing, safety, access, and how Opus 5 compares to Fable 5, Sonnet 5, and the competition.

AI Agent Security

OpenAI's Own AI Agent Broke Containment and Hacked Hugging Face: The AI Agent Security Wake-Up Call

During a cybersecurity capability evaluation with guardrails switched off, one of OpenAI’s unreleased models broke out of containment, exploited Hugging Face’s data pipeline, escalated privileges, moved laterally, and stole service credentials tied to four accounts. Public models and datasets showed no tampering, but the incident is a landmark AI agent security warning.

Gemini 4

Gemini 4 Training Has Begun: What Google Confirmed—and What It Did Not

Google has started what it calls its most ambitious pre-training run yet for Gemini 4, but has not announced a release date, model lineup, benchmarks, pricing, context window, or public preview.

Gemini AI

Gemini 3.6 Flash Launch: Price, Benchmarks, API & Flash-Lite

Google’s July 21 Gemini update adds a more token-efficient 3.6 Flash, a 350-token-per-second 3.5 Flash-Lite, and the restricted Gemini 3.5 Flash Cyber model for CodeMender.

Claude AI

Claude Cowork Can Record Your Workflow and Turn It Into a Skill

Anthropic’s new Record a skill command lets paid Claude users demonstrate a task on screen, explain decisions aloud, and convert that demonstration into a reusable Cowork workflow.

Qwen 3.8

Qwen 3.8: Preview Access, Specs, Pricing & Benchmarks

A source-checked guide to Qwen3.8-Max-Preview covering release status, 2.4T scale, context and output limits, reasoning controls, Qwen Cloud pricing, early coding evidence, and the promised open-weight release.

Kimi K3

Kimi K3 Sold Out: Why Moonshot Paused New Subscriptions

Moonshot AI paused new Kimi subscriptions after Kimi K3 demand pushed GPU capacity close to the limit. Here is what sold out means, what existing users keep, and how Kimi K3 compares with Claude Opus 4.8 and GPT-5.6 Sol.

Codex Merged With the ChatGPT App: What Changed, Access & What's New OpenAI 9 min
OpenAI

Codex Merged With the ChatGPT App: What Changed, Access & What's New

OpenAI says the Codex app is merging with the new ChatGPT desktop app. Codex remains a dedicated coding experience, now alongside Chat and Work in one desktop app.

GPT-5.6 Luna: Fast Low-Cost Model, Pricing, Model ID & Use Cases OpenAI 10 min
OpenAI

GPT-5.6 Luna: Fast Low-Cost Model, Pricing, Model ID & Use Cases

A practical guide to GPT-5.6 Luna, OpenAI’s fast low-cost GPT-5.6 tier for high-volume classification, extraction, drafts, routing, and first-pass automation.