Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Samstag, 8. August 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech22 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
2 / 22
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Samstag, 8. August 2026
AI developer tools · What shipped

Qwen3.8-Max leads; repo benchmarks cap at 55%; Muse Code v1.2 pricing

1 Min. Lesezeit

Qwen 3.8-Max flagship

Alibaba's 2.4T flagship just shipped with serious agentic chops.

Qwen 3.8-Max hit August 7 with 95 billion active parameters and 67% on SWE-Pro, a code completion benchmark that mirrors real shipping tasks [Source: Pat McGuinness]. The model scored 86% on OSWorld-Verified agentic tasks—repository-scale work—and API pricing sits at $2/$6 per million tokens, a middle tier between budget and frontier. This lands two weeks after Qwen3.7 Max dominated the competitive coding leaderboard at 91.6%.

Watch whether the agentic benchmark slice (OSWorld) becomes the new signal for shipping quality.

VIBE-Pro repo benchmark

End-to-end project delivery still maxes out around 55%.

MiniMax M2.7 leads the August VIBE-Pro snapshot at 55.6%, with 23 model releases chasing the leaderboard in the last month [Source: BenchLM]. VIBE-Pro tests whether models can complete substantial product requirements across web, mobile, and simulation tasks—not isolated code snippets. The 55% ceiling signals how much harder full-project delivery is compared to single-file generation, where vision-language code benchmarks already hit 98.8%.

The gap between snippet and system tells you where the real bottleneck still lives.

Muse Spark 1.2 pricing

Meta's terminal agent now underbids the tier below frontier.

Muse Spark 1.2 runs at $0.10/$0.20 per million tokens with contributor tier pricing roughly 12x cheaper on input than standard tier [Source: Patrick McGuinness]. Yesterday's coverage showed the agent carries persistent background workers and replay-safe crash recovery; today's move is pricing alignment below mid-tier alternatives. Muse Code (beta) shipped August 5 on macOS and Linux; the harness-aware coupling means performance degrades outside Meta's environment.

Other vendors are now forced to choose: couple performance tightly or accept the price penalty.

Competitive coding benchmarks

LiveCodeBench v6 now has 16 models; Sakana Fugu-Ultra leads at 93.2%.

August 7 update shows Sakana Fugu at 92.9% and Kimi K2.6 at 89.6% on the v6 named release slice [Source: BenchLM]. LiveCodeBench v6 is kept separate from the rolling leaderboard to prevent mixing named releases with continuous windows; the 21.2-point spread in the top-10 range shows clustering is tightening. Rolling LiveCodeBench still has Qwen3.7 Max at 91.6%, unchanged from three days ago.

Watch for the first model to break 94% on competitive programming tasks.

Quellen
VIBE-Pro Leaderboard & Scores — August 2026 | BenchLM.ai
VIBE-Pro Leaderboard & Scores — August 2026 | BenchLM.ai
10 hours ago ... Benchmark profile. VIBE-Pro. A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.
benchlm.ai
KI-Zusammenfassung

VIBE-Pro is a repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks, with data verified as of August 7, 2026. MiniMax M2.7 leads the public benchmark snapshot with a score of 55.6%, with 23 confirmed releases in the last 30 days. The benchmark tests whether models can complete substantial product requirements through end-to-end software delivery rather than single-file snippets, and BenchLM refreshes the benchmark quarterly.

Quelle öffnen
LiveCodeBench Leaderboard (August 2026): Qwen3.7 Max Leads ...
LiveCodeBench Leaderboard (August 2026): Qwen3.7 Max Leads ...
9 hours ago ... The official suite evaluates code generation, code execution, test-output ... RadarModel updatesRelease timelineAI RaceLLM pricingPrice vs performanceLLM speed ...
benchlm.ai
KI-Zusammenfassung

Qwen3.7 Max leads the LiveCodeBench leaderboard as of August 7, 2026 with a 91.6% score, followed by Qwen3.7 Plus at 89.6% and GLM-4.7 at 84.9% across six tracked models. LiveCodeBench is a continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems, evaluating code generation, execution, test-output prediction, and self-repair on competitive programming tasks. The benchmark contributes 38% of the coding category score in BenchLM's overall model evaluation framework, with results verified and updated rolling through 2026.

Quelle öffnen
LiveCodeBench v6 Leaderboard & Scores - Benchmarks - BenchLM.ai
LiveCodeBench v6 Leaderboard & Scores - Benchmarks - BenchLM.ai
10 hours ago ... The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch ...
benchlm.ai
KI-Zusammenfassung

LiveCodeBench v6, a named release benchmark for code generation, was updated August 7, 2026, with 16 AI models evaluated on competitive programming tasks. Sakana Fugu-Ultra leads at 93.2%, followed by Sakana Fugu at 92.9% and Kimi K2.6 at 89.6%. The benchmark measures contest-style code generation across a 21.2-point spread in the top-10 range and carries 20% weight in BenchLM.ai's overall scoring system, though it is currently displayed for reference only and excluded from the scoring formula.

Quelle öffnen
AI Week in Review 26.08.07 - by Patrick McGuinness
AI Week in Review 26.08.07 - by Patrick McGuinness
21 hours ago ... Muse Spark 1.2 was co-trained with Muse Code to optimize performance for long-horizon coding tasks, including whole-repository generation, debugging, and ...
patmcguinness.substack.com
KI-Zusammenfassung

Alibaba released Qwen 3.8-Max, a 2.4T parameter flagship model with 95 billion active parameters achieving 67% on SWE-Pro coding benchmark and 86% on OSWorld-Verified agentic tasks, with API pricing at $2/$6 per million tokens. Meta released Muse Spark 1.2, a coding-focused update featuring a 1M token context window and scoring 82.9% on Terminal-Bench 2.1 and 59.3% on Deep-SWE, priced at $1.25/$4.25 per million tokens; Meta also beta released Muse Code, a terminal-based coding agent for repository-scale engineering with multi-agent support. Minimax launched Minimax H3, an open-weights omni-modal model for audio-video generation achieving the number two spot on Arena for image-to-video, priced at $0.13 per second for 2K generation. Nvidia released Alpamayo 2 Super, an open 34B parameter reasoning vision-language-action model for autonomous vehicle development combining the Cosmos 3 Super Reasoner with a diffusion-based Action Expert. AWS announced general availability of Web Search on Amazon Bedrock for grounding foundation model responses in current web knowledge and released an automated web insight extraction solution using Amazon Bedrock AgentCore Browser. OpenAI solved ten major open mathematics problems using an internal version of its forthcoming Astra AI model, representing a transition to AI systems contributing genuinely new mathematics. Google expanded Ask Maps with agentic AI assistant tools for restaurant identification and food orders through Toast, Square, and Uber Eats, plus hotel price comparison and event ticket finding. AWS also released an event-driven architecture solution utilizing Amazon Bedrock and OpenSearch Serverless for AI-powered summaries from JavaScript-heavy web pages and RSS feeds.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10