news
AI Daily|Anthropic Drops Claude Opus 5.5, OpenAI Unveils GPT-6 Sol & Luna in Major Price War Model Releases & Updates Claude Opus 5.5 — Anthropic TL;DR: Anthropic has introduced Claude Opus 5.5, the inaugural model in the Claude 5.5 series, bringing Mythos-tier capabilities to everyday workloads at 40% lower total cost. Key Highlights: Delivers performance comparable to Claude Fable 5.1 across agentic coding, computer use, and complex knowledge synthesis, while generating output over 30% faster than Opus 5. Slashes API pricing to $4/M input and $20/M output (a 20% base reduction), with prompt cache reads reduced by 60% down to $0.20/M tokens. Transitions exclusively to adaptive thinking and retires forced tool use in favor of organic steerability, alongside a more concise, front-loaded communication style. Subscription tiers (Pro, Max, Team) receive increased 5-hour rate limits and a bankable on-demand limit reset. Specs: Frontier Flagship / 1M Context Window / Tops Artificial Analysis Intelligence Index (Score: 58), 66.4% on Terminal-Bench 4.0 Links: Anthropic Blog GPT-6 Sol & GPT-6 Luna — OpenAI TL;DR: OpenAI has launched GPT-6 Sol and GPT-6 Luna, bringing the architectural and alignment breakthroughs of its flagship GPT-6 Astra to high-throughput, cost-sensitive production tiers. Key Highlights: Cuts API token pricing by 50% relative to previous GPT-5.6 promotional tiers: GPT-6 Sol is priced at $2/M input and $10/M output, while GPT-6 Luna drops to just $0.10/M input and $0.50/M output. Sol matches or outperforms Claude Opus 5 on AutomationBench (33.2%) at roughly 9% of the cost, while Luna reaches 66.6% on DeepSWE v1.1 for 93–96% lower task expenditure. Adopts Astra’s concise, low-jargon communication style and integrates upgraded prompt caching with up to 90% savings on cache reads. Available immediately across the API, ChatGPT Work, Codex, and integrated platforms like Vercel AI Gateway and OpenRouter. Specs: Mid & Lightweight Frontier Checkpoints / Proprietary API / DeepSWE v1.1: 68.8% (Sol), 66.6% (Luna) Links: OpenAI Announcement Grok 4.7 — xAI / SpaceXAI TL;DR: xAI has rolled out Grok 4.7, boosting model capacity to 2.1 trillion parameters with enhanced multi-hour reasoning and agent execution at unchanged pricing. Key Highlights: Retains the existing pricing of $2/M input and $6/M output across a 500k context window while delivering significantly higher tokens-per-second generation. Jumps from 40.4% to 46.3% on CursorBench 4.0 and achieves 64% on EEBench, demonstrating substantial gains in open-world scaffolding and software debugging. Features adjustable reasoning effort modes (High / xHigh) and day-0 availability on Cursor, Grok Build, and the xAI API. Specs: 2.1T Parameters / Proprietary API & Hosted / CursorBench: 46.3%, EEBench: 64% Links: xAI Announcement MiMo-V2.6 (Pro & Flash) — Xiaomi TL;DR: Xiaomi has open-sourced the MiMo-V2.6 family, establishing a new open-weights Pareto frontier for intelligence-to-cost via large-scale multi-task reinforcement learning. Key Highlights: Employs “MixRL” co-training across 750,000 multi-turn trajectory rollouts spanning coding, cybersecurity, and tool usage. MiMo-V2.6-Pro matches top proprietary tiers with a score of 46 on the Artificial Analysis Intelligence Index, while keeping API rates at $0.435/M input and $0.87/M output (Flash: $0.14/M in, $0.28/M out). Released under the MIT license alongside 7,000+ verifiable RL environments, a full training framework, and a distilled MiMo-V2.6-Distill-Qwen-9B checkpoint. Specs: 1.02T Total / 42B Active MoE (Pro) & Dense 9B Distill / Open Weights / MIT License Links: MiMo Official Release | Hugging Face Repository Product Releases & Updates JetBrains Air — JetBrains What’s New: JetBrains has unified its agentic ecosystem into JetBrains Air, a multi-surface developer platform designed to govern, execute, and monitor autonomous coding agents inside and beyond traditional IDEs. It introduces shared cross-agent context, managed cloud runtimes, automated CI integration, and fine-grained AI token cost governance for enterprise engineering teams. Who It’s For: Software engineering teams, devops leads, and engineering managers scaling agentic development pipelines. Try It: JetBrains Air Overview Worker Previews — Cloudflare What’s New: Cloudflare has launched Worker Previews, providing ephemeral, production-grade staging environments for every Git branch or agent pull request. Each preview receives its own isolated bindings, variables, secrets, Durable Objects, and telemetry, enabling autonomous coding agents to battle-test infrastructure changes without manual orchestration or staging collisions. Who It’s For: Cloud backend engineers, full-stack builders, and developers deploying autonomous DevOps agents. Try It: Cloudflare Blog Advanced Prompt Caching & Diagnostics — OpenAI What’s New: OpenAI has overhauled prompt caching for the GPT-6 ecosystem. Developers can now define explicit cache breakpoints, adjust reasoning effort or toolsets mid-thread without cache invalidation, and prewarm shared enterprise context to minimize time-to-first-token. A dedicated Prompt Caching Dashboard and Diagnostics API provide real-time token reuse rates and cache eviction traces. Who It’s For: Backend AI architects and platform engineers optimizing long-running agents and retrieval systems. Try It: OpenAI Prompt Caching Guide Scribe v2 Medical — ElevenLabs What’s New: ElevenLabs released Scribe v2 Medical, a HIPAA-eligible speech-to-text model specifically fine-tuned for clinical interactions and complex medical terminology. It delivers an 18% lower word error rate (WER) across anatomy, pharmacology, and pathology terms, paired with a strict Zero Retention Mode that wipes audio and transcript data immediately upon request completion. Who It’s For: Healthcare providers, telehealth platforms, and clinical transcription software developers. Try It: View post on X by @ElevenLabs 在 X 上查看 @ElevenLabs 的貼文 ↗
Sep 23, 2026
Read →