New AI tools, tracked the day they drop
Every new model, tool and agent — sourced facts, honest first looks, and a rating only once we've gone hands-on. Never an invented score.
Latest drops
The radarNano Banana 2 Lite
Google's fastest and cheapest Gemini image model. I ran it: a text-heavy poster came back in 4.3 seconds with both lines of type rendered perfectly — the failure mode most image models still trip on.
Claude Sonnet 5
Anthropic's Sonnet-tier model — 'the best combination of speed and intelligence.' Near-frontier on coding and agentic work at roughly a third of Opus's price, with a 1M-token context and adaptive thinking on by default.
Claude Mythos
Claude Mythos is Anthropic's cybersecurity-focused, Mythos-class model — built to autonomously find and fix software vulnerabilities. Its rollout is restricted; a safeguarded public sibling, Fable 5, carries Mythos-class capability to the broader public.
Highlights
1 new this monthBy type
Every AI we've reviewed, by what it is
By access
Open weights, free, freemium or paid
By modality
What these models take in and put out
The index
Every AI we've reviewed — compare at a glance. Sort by type, category, pricing or date.
| # | access | Key factheadline spec | launch | we verified | ||||
|---|---|---|---|---|---|---|---|---|
| 01 | Nano Banana 2 Lite Google | model | image | freemium | ~4s · $0.034/img | 4w ago | 2026-07-19 | Review |
| 02 | Claude Sonnet 5 Anthropic | model | general LLM | paid | 1M context · $2 / $10 per MTok (intro) | 4w ago | 2026-07-24 | Review |
| 03 | Claude Mythos Anthropic | model | security | paid | Restricted release · cybersecurity | 1mo ago | 2026-07-08 | Review |
| 04 | GPT-5.6 OpenAI | model | general | paid | 3 tiers · Sol/Terra/Luna | 1mo ago | 2026-07-21 | Review |
| 05 | Qwen-AgentWorld Alibaba (Qwen Team) | model | agents | open source | 7 envs · 35B/397B | 1mo ago | 2026-07-08 | Review |
| 06 | Sakana Fugu Sakana AI | agent | agents | paid | Limited preview · orchestrator | 1mo ago | 2026-07-08 | Review |
| 07 | GLM-5.2 Zhipu AI (Z.ai) | model | coding | open source | 753B MoE · ~40B active | 1mo ago | 2026-06-26 | Review |
| 08 | VibeThinker-3B Weibo AI | model | reasoning | open source | 3B dense · AIME26 94.3 | 1mo ago | 2026-06-26 | Review |
| 09 | Gemma 4 Google DeepMind | model | multimodal | open source | 128K–256K ctx · 5 sizes | 1mo ago | 2026-06-29 | Review |
| 10 | MiniMax M3 MiniMax | model | multimodal | paid | 428B MoE · 1M ctx | 1mo ago | 2026-07-08 | Review |
| 11 | Claude Opus 4.8 Anthropic | model | frontier LLM | paid | 1M context · $5 / $25 per MTok | 2mo ago | 2026-07-24 | Review |
| 12 | Kling 3.0 Kuaishou | model | video | paid | Up to 15s · native audio | 5mo ago | 2026-06-26 | Review |
| 13 | Kortex for NotebookLM Yaksh Gandhi | tool | knowledge | freemium | 100k users · 4.8★ | 7mo ago | 2026-06-26 | Review |
| 14 | Pyramid Flow Peking University · Kuaishou | model | video | open source | 768p · 24fps · 10s | 1y ago | 2026-06-26 | Review |
| 15 | OpenHands All-Hands-AI | agent | agents | open source | 78.4k GitHub stars | 2y ago | 2026-06-26 | Review |

Nano Banana 2 Lite
Google's fastest and cheapest Gemini image model. I ran it: a text-heavy poster came back in 4.3 seconds with both lines of type rendered perfectly — the failure mode most image models still trip on.
Claude Sonnet 5
Anthropic's Sonnet-tier model — 'the best combination of speed and intelligence.' Near-frontier on coding and agentic work at roughly a third of Opus's price, with a 1M-token context and adaptive thinking on by default.
Claude Mythos
Claude Mythos is Anthropic's cybersecurity-focused, Mythos-class model — built to autonomously find and fix software vulnerabilities. Its rollout is restricted; a safeguarded public sibling, Fable 5, carries Mythos-class capability to the broader public.
GPT-5.6
OpenAI's next flagship after GPT-5 — announced June 26, 2026 as a preview and generally available since July 9, 2026 in three tiers (Sol, Terra, Luna) with published per-token pricing.
Qwen-AgentWorld
Alibaba's Qwen team open-sources AgentWorld — a language world model that simulates seven agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) and is trained to predict how each environment responds to an action. Released June 24, 2026 under Apache 2.0.
Sakana Fugu
Sakana AI's multi-agent orchestration system — Fugu routes complex tasks across a network of specialized agents, all through a single model API endpoint.
GLM-5.2
Zhipu's open-weight coding flagship: a 753B mixture-of-experts model with a 1M-token context and MIT weights, claiming to edge past GPT-5.5 on coding benchmarks.
VibeThinker-3B
A 3-billion-parameter open reasoning model that claims to match systems hundreds of times its size on math and code — and has the AI world arguing about whether the benchmarks are real.
Gemma 4
Google DeepMind's fourth-generation open-weight model family — five sizes from 2B to 31B, Apache 2.0 licensed, with the 12B Unified variant accepting text, image, audio, and video in a single encoder-free architecture.
MiniMax M3
MiniMax's third-generation flagship — M3 is an open-weight 428B-parameter MoE (~23B active per token) using MiniMax Sparse Attention (MSA), with a 1M-token context window and native multimodality (text, image, and video input). It's positioned as the first open-weight model to combine frontier coding, a 1M context, and native multimodality — and can even operate a desktop computer.
Claude Opus 4.8
Anthropic's most capable Opus-tier model, built for complex agentic coding and enterprise work — a 1M-token context, adaptive thinking, and 128K max output, hosted only via the Claude API.

Kling 3.0
Kuaishou's flagship hosted video model — cinematic text- and image-to-video with native audio. I ran it: clearly a tier above the open models, with the usual fast-motion artifacts.

Kortex for NotebookLM
A Chrome extension that bolts import/export, folders, a prompt library and bulk tools onto Google NotebookLM — the power-user layer NotebookLM doesn't ship with.

Pyramid Flow
An open-source, MIT-licensed text- and image-to-video model that makes 768p, 24fps, 10-second clips — and runs on a single consumer GPU with offloading.
OpenHands
An open, self-hosted control center for autonomous coding agents — own the infrastructure, bring your own model, and run agents that edit files and ship code.
Facts are sourced. Opinions are owned. Nothing is faked.
Every spec and price is traced to the vendor or provider API and dated. We only score what we've gone hands-on with — otherwise it's a labeled first look, never an invented number.