Daily|Anthropic Drops
Drops Claude Opus
5.5, OpenAI Unvei
Sol & Luna
Work, Codex, and
like Vercel AI Ga
news
AI Daily|Anthropic Drops Claude Opus 5.5, OpenAI Unveils GPT-6 Sol & Luna in Major Price War
2026-09-23
AI Daily|Anthropic Drops Claude Opus 5.5, OpenAI Unveils GPT-6 Sol & Luna in Major Price War
Model Releases & Updates
Claude Opus 5.5 — Anthropic
- TL;DR: Anthropic has introduced Claude Opus 5.5, the inaugural model in the Claude 5.5 series, bringing Mythos-tier capabilities to everyday workloads at 40% lower total cost.
- Key Highlights:
- Delivers performance comparable to Claude Fable 5.1 across agentic coding, computer use, and complex knowledge synthesis, while generating output over 30% faster than Opus 5.
- Slashes API pricing to $4/M input and $20/M output (a 20% base reduction), with prompt cache reads reduced by 60% down to $0.20/M tokens.
- Transitions exclusively to adaptive thinking and retires forced tool use in favor of organic steerability, alongside a more concise, front-loaded communication style.
- Subscription tiers (Pro, Max, Team) receive increased 5-hour rate limits and a bankable on-demand limit reset.
- Specs: Frontier Flagship / 1M Context Window / Tops Artificial Analysis Intelligence Index (Score: 58), 66.4% on Terminal-Bench 4.0
- Links: Anthropic Blog
GPT-6 Sol & GPT-6 Luna — OpenAI
- TL;DR: OpenAI has launched GPT-6 Sol and GPT-6 Luna, bringing the architectural and alignment breakthroughs of its flagship GPT-6 Astra to high-throughput, cost-sensitive production tiers.
- Key Highlights:
- Cuts API token pricing by 50% relative to previous GPT-5.6 promotional tiers: GPT-6 Sol is priced at $2/M input and $10/M output, while GPT-6 Luna drops to just $0.10/M input and $0.50/M output.
- Sol matches or outperforms Claude Opus 5 on AutomationBench (33.2%) at roughly 9% of the cost, while Luna reaches 66.6% on DeepSWE v1.1 for 93–96% lower task expenditure.
- Adopts Astra’s concise, low-jargon communication style and integrates upgraded prompt caching with up to 90% savings on cache reads.
- Available immediately across the API, ChatGPT Work, Codex, and integrated platforms like Vercel AI Gateway and OpenRouter.
- Specs: Mid & Lightweight Frontier Checkpoints / Proprietary API / DeepSWE v1.1: 68.8% (Sol), 66.6% (Luna)
- Links: OpenAI Announcement
Grok 4.7 — xAI / SpaceXAI
- TL;DR: xAI has rolled out Grok 4.7, boosting model capacity to 2.1 trillion parameters with enhanced multi-hour reasoning and agent execution at unchanged pricing.
- Key Highlights:
- Retains the existing pricing of $2/M input and $6/M output across a 500k context window while delivering significantly higher tokens-per-second generation.
- Jumps from 40.4% to 46.3% on CursorBench 4.0 and achieves 64% on EEBench, demonstrating substantial gains in open-world scaffolding and software debugging.
- Features adjustable reasoning effort modes (
High/xHigh) and day-0 availability on Cursor, Grok Build, and the xAI API.
- Specs: 2.1T Parameters / Proprietary API & Hosted / CursorBench: 46.3%, EEBench: 64%
- Links: xAI Announcement
MiMo-V2.6 (Pro & Flash) — Xiaomi
- TL;DR: Xiaomi has open-sourced the MiMo-V2.6 family, establishing a new open-weights Pareto frontier for intelligence-to-cost via large-scale multi-task reinforcement learning.
- Key Highlights:
- Employs “MixRL” co-training across 750,000 multi-turn trajectory rollouts spanning coding, cybersecurity, and tool usage.
- MiMo-V2.6-Pro matches top proprietary tiers with a score of 46 on the Artificial Analysis Intelligence Index, while keeping API rates at $0.435/M input and $0.87/M output (Flash: $0.14/M in, $0.28/M out).
- Released under the MIT license alongside 7,000+ verifiable RL environments, a full training framework, and a distilled
MiMo-V2.6-Distill-Qwen-9Bcheckpoint.
- Specs: 1.02T Total / 42B Active MoE (Pro) & Dense 9B Distill / Open Weights / MIT License
- Links: MiMo Official Release | Hugging Face Repository
Product Releases & Updates
JetBrains Air — JetBrains
- What’s New: JetBrains has unified its agentic ecosystem into JetBrains Air, a multi-surface developer platform designed to govern, execute, and monitor autonomous coding agents inside and beyond traditional IDEs. It introduces shared cross-agent context, managed cloud runtimes, automated CI integration, and fine-grained AI token cost governance for enterprise engineering teams.
- Who It’s For: Software engineering teams, devops leads, and engineering managers scaling agentic development pipelines.
- Try It: JetBrains Air Overview
Worker Previews — Cloudflare
- What’s New: Cloudflare has launched Worker Previews, providing ephemeral, production-grade staging environments for every Git branch or agent pull request. Each preview receives its own isolated bindings, variables, secrets, Durable Objects, and telemetry, enabling autonomous coding agents to battle-test infrastructure changes without manual orchestration or staging collisions.
- Who It’s For: Cloud backend engineers, full-stack builders, and developers deploying autonomous DevOps agents.
- Try It: Cloudflare Blog
Advanced Prompt Caching & Diagnostics — OpenAI
- What’s New: OpenAI has overhauled prompt caching for the GPT-6 ecosystem. Developers can now define explicit cache breakpoints, adjust reasoning effort or toolsets mid-thread without cache invalidation, and prewarm shared enterprise context to minimize time-to-first-token. A dedicated Prompt Caching Dashboard and Diagnostics API provide real-time token reuse rates and cache eviction traces.
- Who It’s For: Backend AI architects and platform engineers optimizing long-running agents and retrieval systems.
- Try It: OpenAI Prompt Caching Guide
Scribe v2 Medical — ElevenLabs
- What’s New: ElevenLabs released Scribe v2 Medical, a HIPAA-eligible speech-to-text model specifically fine-tuned for clinical interactions and complex medical terminology. It delivers an 18% lower word error rate (WER) across anatomy, pharmacology, and pathology terms, paired with a strict Zero Retention Mode that wipes audio and transcript data immediately upon request completion.
- Who It’s For: Healthcare providers, telehealth platforms, and clinical transcription software developers.
- Try It:
Industry News
Firecrawl Raises $75M Series B and Launches Alexandria Data Exchange
- What Happened: Firecrawl announced a $75 million Series B funding round led by Smash Capital and unveiled Alexandria, a structured data marketplace designed to feed real-time knowledge into agentic search systems. Alexandria aggregates live web indices, academic literature, and government archives, enabling data publishers and creators to earn micro-royalties when AI agents query their content.
- Why It Matters: As frontier models increasingly depend on real-time web interaction, data extraction infrastructure has become high-value real estate. Alexandria creates a formal economic layer for agentic data consumption, addressing long-standing publisher monetization tensions.
- Source: Firecrawl Alexandria Announcement
Snorkel AI Secures $350M Series E at $3.5B Valuation
- What Happened: Snorkel AI closed a $350 million Series E round co-led by Insight Partners and S32, tripling its valuation to $3.5 billion. The company reported $375 million in annualized revenue—an 18x surge over the past 12 months—driven by intense demand from frontier AI labs for high-grade programmatic data labeling and custom reinforcement learning environments.
- Why It Matters: The shift from generic pre-training data to verifiable, domain-specific post-training trajectories and RL environments has created massive enterprise moats for programmatic data platforms.
- Source: TechCrunch Coverage
a16z Incubates The Horowitz Andreessen Academy (HAA)
- What Happened: Venture firm Andreessen Horowitz announced the launch of The Horowitz Andreessen Academy, an in-person, project-based educational institution in San Francisco set to open in Fall 2027. Backed by founding industry partners including OpenAI, Anthropic, NVIDIA, Meta, Google, Anduril, Stripe, and Palantir, the academy replaces traditional lectures with immersive AI tooling, real-world engineering sprints, and corporate apprenticeships.
- Why It Matters: Frontier AI leaders are actively bypassing legacy university curricula to construct custom talent pipelines tailored for the AI-native economy.
- Source:
Announcing The Horowitz Andreessen Academy.
— a16z (@a16z) September 22, 2026
The internet made it possible to know anything. AI makes it possible to do anything.
The best assignments now are problems so hard they cannot be solved without AI. The Academy courses follow the same logic and will be taught by… https://t.co/fsf75UYdmA
US Treasury & White House Push Back on AI Liability Waivers
- What Happened: In a national broadcast, US Treasury Secretary Scott Bessent clarified that the federal government will not grant blanket liability protections or antitrust safe harbors to frontier AI labs. Stressing that existing product liability, cybersecurity, and contract laws apply equally to AI companies, Bessent also warned that upcoming tech IPOs will be required to disclose concrete risk factors regarding autonomous agent safety.
- Why It Matters: This stance directly counters recent calls by major AI labs for state-supervised safety pacts, placing full legal responsibility for software failures and autonomous agent actions squarely on model builders and enterprise deployers.
- Source: Semafor & Financial Reports
Research Papers
DualSQL: Robust Text-to-SQL via Multi-Agent Reinforcement Learning — Google Research
- Motivation: Traditional single-model Text-to-SQL systems often suffer from catastrophic alignment drift when attempting to balance schema linking and syntactically complex query generation across large enterprise databases.
- Key Innovation: DualSQL splits the generation process into two specialized agents—a schema linker and a query writer—trained end-to-end within a single set of shared weights via Multi-Agent RL. It incorporates an execution-level verification reward and rollout stability guardrails to prevent agent collapse.
- Results: DualSQL-4B matches prior 7B baselines at 68.0% execution accuracy on BIRD dev with only 3,755 training samples, while DualSQL-8B scores 71.1%, outperforming previous 32B single-model solutions.
- Paper: arXiv:2609.18135
Self-Correction via Hint-Guided Self-Distillation for Computer Agents — Perplexity
- Motivation: Computer-use and web-browsing agents frequently fail at cascading tool-calling steps due to compounding environmental errors and rigid error-recovery policies.
- Key Innovation: Researchers introduced a hint-guided self-distillation post-training framework where models iteratively review their own failed trajectory rollouts with targeted corrective signals, distilling the learned recoveries back into the base policy without requiring hints at inference time.
- Results: In live production A/B evaluation, post-trained checkpoints reduced unprompted tool-call failures by 21.2% while boosting zero-shot failure recovery from 75.1% to 93.7%.
- Paper: Perplexity Research Blog
Other Highlights
Figure Helix 2.5: Zero-Shot Autonomous Humanoids in 30 Real Homes
- Overview: Robotics venture Figure released Helix 2.5 alongside 4 hours of unedited evaluation footage showcasing humanoid robots performing multi-step domestic tasks across 30 real SF Bay Area rental homes with zero prior scene training or environment-specific calibration.
- Link: Figure Technical Breakdown
DolphinBench: Action-Based Evaluation for Long-Term Agent Memory
- Overview: Mem0 launched DolphinBench, a public benchmark mapping the Pareto frontier of agent memory systems. Unlike static multiple-choice tests, DolphinBench grades downstream tool actions across lengthy, noisy interaction histories while tracking execution latency, precision, and token cost on a single unified leaderboard.
- Link: DolphinBench Platform

