AI Daily | Kolibri Open-Source MoE Model, Perplexity Visualize Feature, and Meta RankEvolve Multi-Agent Framework
Model Releases & Updates
Kolibri — Aleph Alpha — Aleph Alpha
- Bottom line:Aleph Alpha has released Kolibri, a 78.1B parameter open-weight German-English bilingual MoE model featuring a 1M token context window under the Apache 2.0 license.
- Architecture:Employs a Mixture-of-Experts architecture with 78.1B total parameters and 3.46B active parameters per token.
- Context & Openness:Supports up to 1 million tokens of context and is fully open-weighted under the Apache 2.0 license.
- Source:Tech Report
- Source:Hugging Face
Sopro V2 Turbo 2610 — Samuel Vitorino — Samuel Vitorino / LocalLLaMA
- Bottom line:An interim update to the lightweight Sopro text-to-speech model focusing on cleaner voice cloning and reduced roughness while maintaining a 120M parameter size.
- Performance:Achieves ~300ms to first audio on laptop CPUs with a 120M parameter architecture.
- Languages:Supports English, European Portuguese, French, and German under an Apache 2.0 license.
- Source:GitHub Repository
- Source:Hugging Face Weights
Product Releases & Updates
Visualize Feature in Perplexity Computer — Perplexity — Perplexity
- Bottom line:Perplexity introduced a new Visualize capability in Computer, generating inline interactive components and animations directly within user chat responses.
- Functionality:Users can prompt queries starting with ‘Visualize’ to receive dynamic educational widgets and animations.
- Recommendation:Recommended for use under Standard or High effort modes for optimal visual output.
- Source:
Extract v2.5 — LlamaIndex — LlamaIndex
- Bottom line:LlamaIndex launched Extract v2.5, a series of frontier agents tuned for precise multi-page document extraction and complex tabular parsing.
- Performance:Outperforms Opus 5.5 and GPT-6 Sol on complex long-list and multi-page table extractions while being 30% to 4x cheaper.
- Accuracy Gains:Achieved a jump from 86.1% to 95.5% on complex extraction tasks using the agentic tier.
- Source:Blog Announcement
Industry News
AWS Ends Government NDA Usage for Data Centers — Amazon Web Services — Amazon Web Services
- Bottom line:AWS CEO Matt Garman announced that the company has stopped using non-disclosure agreements when partnering with government agencies on data center construction.
- Policy Shift:Eliminated NDAs for government-linked data center projects to increase transparency amid public scrutiny.
- Resource Defense:Garman noted that direct data center water usage accounts for only 0.5% of US industrial consumption and highlighted $1 billion in community investments.
- Source:TechCrunch Coverage
Cloudflare Launches Next-Gen Git Platform Challenge — Cloudflare — Cloudflare
- Bottom line:Cloudflare announced an open beta for Artifacts and launched a developer challenge to build multi-agent Git platforms using Workers and Artifacts.
- Competition Details:Invites developers to build agent-era Git platforms requiring multi-agent concurrency, running until October 14, 2026.
- Prizes:Offers $25,000 in Cloudflare credits for the first-place winner.
- Source:Cloudflare Blog
Research Papers
RankEvolve: Multi-Agent Auto-Research Framework — Meta — Meta
- Bottom line:Meta introduced RankEvolve, a multi-agent framework where coding models cross-review and refine research code modifications under compilation protocols.
- Methodology:Leverages models like Claude Code and Codex as independent nodes to mutually review and patch code changes.
- Results:Boosted task execution accuracy from 45.8% to 62.5% and improved recommendation model benchmarks.
- Source:
The Sharpening Tax in RL Post-Training — Meta Superintelligence Labs — Meta Superintelligence Labs
- Bottom line:Researchers discovered that reinforcement learning post-training improves pass@1 accuracy but imposes a ‘Sharpening Tax’ that limits test-time scaling performance.
- Finding:Base models with lightweight inference frameworks frequently outperform RL-tuned models when given sufficient sampling budgets.
- Proposed Solution:Introduced PTGS to estimate prompt difficulty and set sampling temperatures to mitigate the scaling tax.
- Source:
ThinkingBox Agent Sandbox & Benchmark — Microsoft & Hugging Face — Microsoft / Hugging Face
- Bottom line:Microsoft and Hugging Face released ThinkingBox, an agent sandbox and benchmark designed to evaluate stateful business workflows using database end-states.
- Scope:Covers 507 stateful business workflows with 20 repeated evaluations per task.
- Execution:Uses executable evaluation based on final database states and side effects, accessible via OpenEnv.
- Source:Hugging Face Blog
Other Highlights
Advocating Default Hard Budget Caps for Metered AI Services — Simon Willison — Simon Willison
- Bottom line:Developer Simon Willison published an essay calling for metered AI services and APIs to provide mandatory default hard budget caps to prevent runaway agent bills.
- Core Argument:Usage-based services should immediately cut off and return errors upon hitting limits rather than relying solely on warning emails.
- Ecosystem Impact:Autonomous coding agents lower the barrier to deploying paid code, escalating the risk of accidental large invoices.
- Source:Blog Post

