AI Daily | Amazon Triples NVIDIA Order, Perplexity Launches Brain, & AWS Unpacks Agent Handoff Tax
Model Releases & Updates
Tenet — Harvey & Fireworks AI
- TL;DR: Harvey and Fireworks AI introduced Tenet, a specialized frontier legal reasoning model post-trained from Kimi K3 using asynchronous reinforcement learning.
- Key Highlights:
- Solves nearly 2x as many complex long-horizon legal analysis tasks on the LAB benchmark compared to base Kimi K3.
- Maintains flat inference costs ($5.92 per LAB task vs. $5.62 for base K3) through custom reward shaping designed to penalize redundant token generation.
- Specs: Post-trained Kimi K3 foundation / Asynchronous RL / API available via Fireworks AI & Harvey
- Links:/ Fireworks Blog
Recently @Harvey introduced Tenet, its first model, trained for long horizon legal work.
— Fireworks (@FireworksAI_HQ) August 26, 2026
Harvey post-trained it from a Kimi K3 base in collaboration with Fireworks using async RL on our Training API.
A thread on promising initial results for both performance and… pic.twitter.com/7JNtuattPz
Apodex 1.1 & Apodex 1.1 Mini (35B) — Apodex
- TL;DR: Apodex released its 1.1 model family alongside open weights for the 35B parameter Mini model, engineered specifically for long-running, autonomous multi-agent deep research.
- Key Highlights:
- Coordinates up to 150 specialized sub-agents running asynchronous file inspection, data scrubbing, code execution, and fact verification into a shared pool.
- Features built-in execution recovery that dynamically amends investigation plans upon encountering runtime code errors or data anomalies rather than restarting tasks.
- Specs: 35B open weights / Multi-agent orchestration framework (FrontierAgent) / MIT License
- Links: Apodex Blog / GitHub
Wan 3.0 Prime — Alibaba Wan
- TL;DR: Alibaba rolled out Wan 3.0 Prime, an accelerated variant of its Wan 3.0 video foundation model delivering 7x faster generation speeds for audio-visual video creation.
- Key Highlights:
- Generates coherent 30-second video clips with native synchronized sound effects and background audio in a single generation pass.
- Supports up to 20 reference images for consistent character and scene conditioning across text-to-video and image-to-video workflows.
- Specs: Single-pass text-to-video & image-to-video / 30-second max duration / Available via Replicate and Pika API Club
- Links:/
Wan 3.0 Prime is now on Replicate.
— Replicate (@replicate) August 26, 2026
The accelerated variant of Wan 3.0 from @Alibaba_Wan for text-to-video, image-to-video, and reference-driven workflows.
Generate clips up to 30 seconds long in one shot with integrated audio-visual generation. pic.twitter.com/dJG4XPIDeeJust when you thought WAN 3.0 couldn't get more exciting...enter WAN 3.0 Prime.
— Pika (@pika_labs) August 26, 2026
Same outstanding quality, 7x faster generation time. https://t.co/g1xKQJqqPa pic.twitter.com/TqfpEWhZ1M
Hy-MT2-1.8B (Sherry & SEQ Quantization) — Tencent Hunyuan
- TL;DR: Tencent Hunyuan open-sourced extreme low-bit quantized builds of its Hy-MT2-1.8B on-device translation model, compressing the footprint down to 440MB with near-zero quality loss.
- Key Highlights:
- Uses 1.25-bit Sherry (Sparse High-Efficiency Ternary Quantization) and 2-bit Structured Elastic Quantization (SEQ) paired with SIMD-optimized x86 kernels.
- Deployed into production across Bilibili live streams for sub-800ms real-time bullet chat translation, outperforming commercial APIs on FLORES-200.
- Specs: 1.8B base / 440MB (1.25-bit) & 574MB (2-bit) / 33 languages / x86 & ARM CPU optimized
- Links: Tencent Hunyuan WeChat
Muse Image — Meta Superintelligence Labs
- TL;DR: Meta Superintelligence Labs released Muse Image, a unified image generation and editing foundation model now accessible on developer gateways and creative platforms.
- Key Highlights:
- Unifies text-to-image synthesis and iterative image-to-image editing within a single checkpoint, eliminating the need to toggle between separate generation and inpainting models.
- Accepts reference conditioning images to steer stylistic textures, palettes, and compositions while executing targeted local modifications from natural language prompts.
- Specs: Unified generative & editing architecture / Available via Vercel AI Gateway (
meta/muse-image-1.0) and Runway ML - Links: Vercel Changelog /
Muse Image from Meta, now on Runway.
— Runway (@runwayml) August 26, 2026
Available today alongside the world's best image and video models. Try it now at the link below. pic.twitter.com/bA2yBNfbEc
Product Releases & Updates
Gemini Live Personal Intelligence & Daily Brief — Google
- What’s New: Google upgraded Gemini Live with Personal Intelligence and Daily Brief, transforming the conversational voice interface into an autonomous task delegate. Gemini Live now securely synthesizes context across Gmail, Google Photos, Google Calendar, and YouTube history, allowing users to talk through and execute cross-app to-dos, inbox triage, and daily scheduling hands-free.
- Who It’s For: Busy professionals, mobile users, and knowledge workers seeking voice-first productivity automation.
- Try It: Google Blog /
Gemini Live is moving beyond conversation to handle complex tasks on your behalf.
— Google Gemini (@GeminiApp) August 26, 2026
With new features in Live like Daily Brief, Gemini Spark, Personal Intelligence, and @Gmail inbox management, you can talk through your day and delegate your to-dos without missing a beat. 🧵
Brain Memory System for Perplexity Computer — Perplexity
- What’s New: Perplexity unveiled Brain, a self-improving persistent memory architecture for Perplexity Computer. Rather than loading massive conversation logs into every prompt, Brain runs offline background “Dream agents” to periodically synthesize past sessions, files, and research notes into a structured wiki, improving answer correctness by 9.3 points and recall by 8.9 points while cutting active token consumption by 15%.
- Who It’s For: Researchers, analysts, and enterprise users managing long-term, multi-session knowledge work.
- Try It:/ Perplexity Blog
Brain is our self-improving memory system for Perplexity Computer. It compiles sessions, files, and sources into a structured knowledge wiki.
— Perplexity (@perplexity_ai) August 26, 2026
New evals build on our initial results, improving correctness by 9.3 points, currentness by 8.0, and recall by 8.9 with 15% fewer tokens. pic.twitter.com/mDMVWt2xzS
Visual Studio Debugger Agent: Test-Driven Investigation — Microsoft
- What’s New: Microsoft expanded the Visual Studio Debugger Agent with an autonomous Test-Driven Investigation workflow. When presented with intermittent bug reports or stack traces, the agent automatically synthesizes isolated unit tests, steps through the live debugger to isolate the root cause, and re-executes tests to verify proposed patches before submission.
- Who It’s For: .NET and C++ developers debugging non-deterministic errors and flaky test suites.
- Try It: Microsoft DevBlog
Grok Voice on LiveKit & Grok Bot Vault — xAI / SpaceX AI & LiveKit
- What’s New: xAI partnered with LiveKit to release an end-to-end voice agent pipeline (Grok STT → Grok 4.3 → Grok TTS) supporting Zero Data Retention (ZDR) on every hop without requiring multiple API keys. Concurrently, Grok Bot introduced an agent credential vaulting system that prevents LLMs from ever seeing raw plaintext API keys or environment variables during tool execution.
- Who It’s For: Voice AI builders, healthcare developers requiring HIPAA-grade compliance, and security-conscious agent engineers.
- Try It:/
Use Grok Voice models in LiveKit to build useful voice agents with full support for ZDR https://t.co/32Ab1mwYVT
— SpaceXAI (@SpaceXAI) August 26, 2026this card allows grok bot to store credentials securely in a vault! when the credential is needed, it goes through a special access path where its never exposed to the model in plain text
— eric zakariasson (@ericzakariasson) August 26, 2026
later when you say "file a bug that login is broken", the bot turns that into a linear call… https://t.co/lCV6EcTdUZ pic.twitter.com/KkzOUNeyFQ
AI Agents & Model Context Protocol Support — JetBrains (DataGrip & Compose Multiplatform)
- What’s New: JetBrains introduced deep AI agent integrations across its developer ecosystem. DataGrip now includes built-in Model Context Protocol (MCP) servers allowing agents like Claude Code and Codex to inspect database schemas, validate dependencies, and run natural-language SQL queries. Simultaneously, Compose Multiplatform 1.12.0 added an MCP server for Compose Hot Reload, enabling coding agents to inspect the live UI semantic tree and verify visual rendering directly.
- Who It’s For: Full-stack developers, database administrators, and UI engineers leveraging agentic coding workflows.
- Try It: JetBrains DataGrip Blog / JetBrains Kotlin Blog
Industry News
Amazon Triples NVIDIA Order with 2M Additional GPUs for AWS
- What Happened: Amazon and NVIDIA significantly expanded their cloud infrastructure partnership, committing to deploy 2 million additional NVIDIA GPUs—including Blackwell Ultra, Rubin, and Rubin Ultra architectures—across AWS data centers throughout 2027 and 2028.
- Why It Matters: Coming just five months after AWS committed to 1 million GPUs, the tens-of-billions deal highlights soaring enterprise demand for next-generation frontier training and agent inference clusters.
- Source: TechCrunch
NVIDIA Reportedly Commits $6B to Poolside Alliance for Open Nemotron Models
- What Happened: The Wall Street Journal reported that NVIDIA is investing $1 billion into AI coding startup Poolside at a $12 billion valuation while allocating an additional $5 billion to license Poolside’s technology and absorb over 100 staff into its Nemotron open-weight model initiative.
- Why It Matters: Signals NVIDIA’s aggressive pivot toward cultivating sovereign, world-class open-weight foundation models to counter Chinese open-source momentum and reduce industry dependence on closed American proprietary labs.
- Source:
据《华尔街日报》报道,Nvidia计划投入60亿美元,打造全球最强的开放权重AI模型之一。
— AI Will (@FinanceYF5) August 26, 2026
Nvidia将获得Poolside技术授权,并把其100多名员工纳入Nemotron项目;同时再向Poolside投资10亿美元,投前估值120亿美元。
目标是挑战DeepSeek、Kimi等中国开放模型,并与OpenAI、Anthropic直接竞争。 pic.twitter.com/gbgXPkP5Vx
Figure AI Launches Project Index with $1B Commitment for Physical AI Data
- What Happened: Robotics unicorn Figure launched Project Index on iOS and Android to crowdsource real-world physical manipulation datasets. The program has already mobilized over 44,000 weekly active contributors across 108 countries to upload 16 million task demonstration videos, backed by a $1 billion compute and data budget over the next 12 months.
- Why It Matters: Physical AI and humanoid foundation models (such as Figure’s Helix) face severe data bottlenecks that text and synthetic simulations cannot resolve; large-scale human behavioral telemetry is becoming the next critical AI moat.
- Source: Figure Project Index via WeChat
OpenAI Outlines Platform Evolution and Jalapeño Custom ASIC Milestones
- What Happened: In an interview with TIME and technical disclosures, OpenAI CEO Sam Altman and engineering leadership detailed the roadmap toward an internal AGI system by late 2026 alongside progress on their custom Jalapeño inference ASIC. Designed with extensive hardware-in-the-loop assistance from internal AI models in just 16 months, the chip achieved over 700–1400 tokens/s per user on reasoning workloads with leading performance-per-watt metrics.
- Why It Matters: Underscores OpenAI’s shift from a pure consumer product company to a full-stack platform provider while aggressively lowering token inference costs through vertical hardware integration.
- Source:/
Sam Altman told TIME that by the end of this year, OpenAI will have an internal system he would call AGI.
— The Rundown AI (@TheRundownAI) August 26, 2026
A few other eye-opening quotes from the interview:
"I expect this will be the first model where the model actually invents new things in a way that matters. That’s a very… pic.twitter.com/iAerBfT1rK/ OpenAI Reportai for chip design is underrated https://t.co/pmCgaGsdNe
— Greg Brockman (@gdb) August 26, 2026
Research Papers
Measuring the Agent Handoff Tax in Multi-Model Cascades — AWS AI Labs
- Motivation: Practical agent systems frequently escalate difficult subtasks to expensive frontier models mid-run to save compute, but the latent performance and monetary overhead of transferring execution state across model families remained unquantified.
- Key Innovation: Evaluated cross-model trajectory handoffs between Claude and GPT series on coding benchmarks, isolating the exact degradation caused by forcing a receiving model to parse another model’s reasoning history.
- Results: Escalating mid-trajectory recovered less than half the quality gap between weak and strong models while introducing a heavy token cost penalty (“handoff tax”); conversely, pruning the weak model’s conversational trajectory significantly restored downstream execution accuracy.
- Paper: ArXiv:2608.24358 /
Great new paper from AWS on agent handoff tax.
— elvis (@omarsar0) August 26, 2026
If you build agents today, you need to understand the so-called handoff tax.
(bookmark it)
Escalating to a stronger model mid-run is usually the resort when a cheap agent stalls.
New work from AWS AI Labs measures how much that… pic.twitter.com/HlDLlw8N2c
Disentangling Harness and Model Variance on SWE-bench — Independent Benchmark Study
- Motivation: Public AI coding leaderboards rank models assuming benchmark scores reflect raw model intelligence, obscuring the impact of execution harnesses, prompt wrappers, and retry heuristics.
- Key Innovation: Constructed a controlled evaluation grid across three frontier models and three standardized harnesses on 100 SWE-bench Verified tasks with identical step budgets, tools, and execution environments.
- Results: Harness-induced variance was 7.8x larger than model-induced variance (swapping harnesses moved scores by up to 13.0 points), causing 6 out of 9 head-to-head model comparisons to flip rank purely based on the harness selected.
- Paper:
Great paper on why agent leaderboard comparisons are hard to trust.
— elvis (@omarsar0) August 26, 2026
It's on the hot topic of how much of an agent benchmark score actually belongs to the harness.
The harness is the layer between the model and the task. It builds the context the model sees, mediates tool calls,… pic.twitter.com/djFsrVaBHs
Other Highlights
Android C2PA Cryptographic Signature Bypass
- Overview: Security researcher David Buchanan published a vulnerability writeup demonstrating that C2PA hardware-backed provenance verification on Android can be forged via StrongBox privilege escalation (CVE-2026-43499), allowing attackers to generate valid cryptographic camera signatures on arbitrary synthesized media.
- Link: David Buchanan Blog
3D Procedural “Tree QR” Generator
- Overview: A creative open-source WebGL tool called Tree QR turns ordinary URLs into interactive 3D digital bonsai trees with scannable QR ground grids, rendering seasonal spring, summer, and autumn themes entirely inside the browser.
- Link: Tree QR App /
这玩意儿真的是艺术品 是谁做的啊 真漂亮真有想象力https://t.co/rmENYVkRUD
— Viking (@vikingmute) August 26, 2026
二维码藏进一棵树的底座里面,每棵树都根据 URL 变化不同,还有春天,夏天和秋天三种不同的风格,点击树真正的二维码就出现了 像个微缩小景观。
这种看起来没什么用的东西为什么会让人这么愉悦呢?这就是美吧… https://t.co/82lGy2mu7l pic.twitter.com/4WvBDBF43q



