Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Sonntag, 16. August 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech35 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
7 / 35
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Sonntag, 16. August 2026
AI developer tools · What shipped

Frontier models flood benchmarks, local agents tighten

1 Min. Lesezeit

Weekly benchmark churn

Six benchmarks updated. Fifteen models shifted rank.

Claude Fable 5 leads the OpenHands Index coding-agent benchmark at 81.0% across issue resolution, frontend work, greenfield development, testing, and information gathering [Source: BenchLM]. On tool calling, Muse Spark 1.1 tops the MCP Atlas leaderboard at 88.1%, ahead of Claude Opus 5 (85.8%) and Kimi K3 (84.2%). React Native Evals show Composer 2 holding 96.1% on mobile app tasks [Source: BenchLM].

The math benchmark still carries the oldest story: Claude Opus 4.6 at 99.8% on AIME25.

Five frontier releases landed

SpaceXAI, Z.ai, Alibaba, and DeepSeek all shipped this week.

Grok 4.6 scores 61 on the Intelligence Index at $2/$6 per million tokens with 69.9% on CursorBench and 61.3% on FrontierCode 1.1 [Source: Patrick McGuinness]. GLM-5.3 outcompetes Kimi K3. Qwen 3.8 Max ships as open weights with 2.4T total parameters and 262K-token context. DeepSeek V4 Pro 0813 scores 87.9% on Terminal-Bench and 62.7% on DeepSWE with native 262K support. Smaller models also moved: Qwen 3.8-27B at 42% on DeepSWE, Muse Glimmer 30B under Apache 2.0, Nemotron 3.5 Lightning 30B supporting one-million-token contexts.

Agent frameworks got real tools: DeepSeek Harness accumulated 23K GitHub stars in days.

Developer tools cut costs, speed up

Three frameworks shipped better bang-per-token this week.

Microsoft released MAI-Code-1.1-Flash integrated into GitHub Copilot with 25% greater token efficiency at one-quarter the cost [Source: Patrick McGuinness]. Writer released Palmyra X6 reducing operating costs and improving speed by over 40%. SpaceXAI shipped Grok Bot for persistent AI agents available on desktop and iOS, pricing half of recent proprietary tiers.

The efficiency gains suggest the pricing war is moving from the model layer into the agent loop.

Quellen
AI Week in Review 26.08.15 - by Patrick McGuinness
AI Week in Review 26.08.15 - by Patrick McGuinness
6 hours ago ... All are significant improvements on intelligence and performance from prior versions released just weeks or months ago. ... Benchmarks for the AI model releases ...
patmcguinness.substack.com
KI-Zusammenfassung

Google released Gemini 3.7 Flash, a multimodal model improving substantially over Gemini 3.6 Flash with gains from 49% to 65% on DeepSWE and from 17% to 30.4% on Automation Bench, achieving an Intelligence Index score of 56 comparable to GPT-5.6 Terra at $0.75/$3.75 per million input/output tokens, half the previous pricing. Five frontier-level models entered the top 20 this week: SpaceXAI released Grok 4.6 scoring 61 on the Intelligence Index at $2/$6 per million tokens with 69.9% on CursorBench 3.2 and 61.3% on FrontierCode 1.1; Z.ai released GLM-5.3 outcompeting Kimi K3; Alibaba published Qwen 3.8 Max as open weights with 2.4T total parameters and 95B active parameters supporting 262,144-token native context; DeepSeek released V4 Pro 0813 with 1.7T total parameters and 49B active parameters scoring 87.9% on Terminal-Bench 2.1 and 62.7% on DeepSWE. Local AI model releases included Qwen 3.8-27B achieving 42% on DeepSWE 1.1 and 73% on Terminal Bench 2.1, Meta's Muse Glimmer 30B under Apache 2.0, Nvidia's Nemotron 3.5 Lightning 30B MoE model supporting one-million-token contexts, and Cohere's North Micro Vision Instruct 2.4B vision-language model. For developer tools, DeepSeek released DeepSeek Harness for agent execution accumulating 23,000 GitHub stars within days, SpaceXAI released Grok Bot for persistent AI agents available on desktop and iOS, Microsoft released MAI-Code-1.1-Flash integrated into GitHub Copilot with 25% greater token efficiency at one-quarter the cost, and Writer released Palmyra X6 reducing operating costs and improving speed by over 40%. Video and image generation updates included Lightricks' LTX-2.5 22B diffusion model for text-to-video and image-to-video supporting multi-shot sequences, Alibaba's Wan-Animate-2 14B for character animation, and MiniMax Music 3 for music generation.

Quelle öffnen
OpenHands Index Leaderboard & Scores - Benchmarks - BenchLM.ai
OpenHands Index Leaderboard & Scores - Benchmarks - BenchLM.ai
21 hours ago ... It is a valuable agentic software-engineering reference, but its rows combine model, SDK version, agent harness, cost, runtime, and per-benchmark result links, ...
benchlm.ai
KI-Zusammenfassung

Claude Fable 5 leads the OpenHands Index coding-agent benchmark with an 81.0% average score across five categories: issue resolution, frontend work, greenfield development, testing, and information gathering. The July 31, 2026 snapshot evaluated 33 AI model variants, with Claude Opus 4.8 (71.9%) and Claude Opus 4.7 Adaptive (69.7%) following in second and third place. OpenHands Index is currently displayed on BenchLM as a reference benchmark but excluded from overall model rankings, with quarterly refresh cadence to track performance changes across real-world software engineering agent tasks.

Quelle öffnen
React Native Evals Leaderboard & Scores - Benchmarks - BenchLM.ai
React Native Evals Leaderboard & Scores - Benchmarks - BenchLM.ai
21 hours ago ... Know when it's worth switching models ... The model to choose, the cheaper alternative, and the release we would wait on. ... Join 2,000+ readers. ... One email each ...
benchlm.ai
KI-Zusammenfassung

Composer 2 leads the React Native Evals benchmark at 96.1%, followed by Composer 2 Fast (94.9%) and GPT-5.4 (85.3%). The open benchmark evaluates AI coding agents on real-world React Native implementation tasks, with 16 models assessed across framework-specific mobile app development work including navigation, animation, and async state management. The benchmark was refreshed August 15, 2026 and operates on a quarterly refresh cadence, currently displayed as reference material with 20% weight in BenchLM's overall scoring system.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10