Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Montag, 14. September 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech66 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
9 / 66
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Montag, 14. September 2026
AI developer tools · What shipped

Fable 5.1 benchmarks confirmed; GPT-6 Astra and Gemini 3.8 ship

1 Min. Lesezeit

Claude Fable 5.1

The benchmark gains hold up under fresh scrutiny.

Fable 5.1's Terminal-Bench 4.0 score of 55.8% and Humanity's Last Exam performance at 65.0% with tools remain the frontier marks [Quelle: Anthropic]. Mythos 5.1, identical internals with 60% fewer safety interventions for vetted security and biotech teams, now has a formal trusted-access program. The $0.25 per-million-token cache-read pricing (75% cheaper than Fable 5) sustains the 45% cost cut for agentic workloads.

Enterprise defaults harden around Fable as competing open weights climb.

GPT-6 Astra models

OpenAI's five-tier lineup covers the full cost-to-capability spectrum.

The max variant scores 53 on Artificial Analysis's Intelligence Index v4.3 (incorporating Terminal-Bench 4.0, SciCode, and Humanity's Last Exam), while low-tier models start at 34 and $0.82 per task [Quelle: Artificial Analysis]. Pricing spans up to 4x across the lineup; output speeds reach 60 tokens per second on max. The tiering forces clearer deployment decisions than prior OpenAI releases.

Watch which tier enterprises standardize on for production agents.

Gemini 3.8 Flash variants

Google's three-tier Flash claims the efficiency crown once again.

The high-intelligence variant scores 41 on the Artificial Analysis Index while delivering 278 tokens per second and 12.81s time-to-first-token [Quelle: Artificial Analysis]. Medium and low variants trade capability for speed and cost, spanning text, image, speech, and video inputs with 1M-token context. The pricing structure forces a sharper tradeoff than Gemini 3.7.

Medium tier may become the standard enterprise choice.

Quellen
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
4 hours ago ... Our most advanced models for coding and knowledge work. Their research capabilities also offer an early glimpse of how AI models will contribute to ...
anthropic.com
KI-Zusammenfassung

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, new models designed for coding and knowledge work with significantly improved capabilities across multiple benchmarks. Fable 5.1 achieves approximately 25% cost reduction for typical workloads and up to 45% for agentic tasks through 75% cheaper cache read pricing ($0.25 per million tokens). The models show substantial performance gains: on Terminal-Bench 4.0, Fable 5.1 scores 55.8% versus Fable 5's 42.0%; on Humanity's Last Exam, it reaches 65.0% with tools versus Fable 5's 63.8%; on CursorBench 3.2.0, it scores 73.4% versus Fable 5's 70.5%. Mythos 5.1 is identical but with reduced safeguards for cybersecurity and life sciences professionals through trusted access programs. The release includes improved safeguards that reduce false positives by 60% in cybersecurity tasks, Enterprise Frontier Safeguards allowing zero-data-retention for enterprise customers, and a new detection API for EU AI Act compliance watermarking.

Quelle öffnen
Gemini 3.8 Flash: Release Intelligence, Performance & Price
Gemini 3.8 Flash: Release Intelligence, Performance & Price
5 hours ago ... The Gemini 3.8 Flash release offers 3 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key ...
artificialanalysis.ai
KI-Zusammenfassung

Google released Gemini 3.8 Flash in September 2026 as a proprietary model with three variants offering different intelligence-performance-cost tradeoffs. The high variant scores 41 on the Artificial Analysis Intelligence Index v4.3 (which incorporates 10 evaluations including AA-Briefcase, AutomationBench-AA, Terminal-Bench v4.0, SciCode, and Humanity's Last Exam), achieves 278 tokens/second output speed with 12.81s latency, and costs $1.24 per Intelligence Index task. The medium variant scores 40 and costs $0.93 per task, while the low variant scores 34. All three support text, image, speech, and video inputs with a 1M token context window and reasoning capabilities.

Quelle öffnen
GPT-6 Astra Models - Intelligence, Performance & Price Comparison
GPT-6 Astra Models - Intelligence, Performance & Price Comparison
3 hours ago ... ... , Video; Inference; Leaderboards; About. AI Trends. Arenas. Premium · Log in. K. All Releases•. OpenAI logo. OpenAI. •. Proprietary release. •. Released ...
artificialanalysis.ai
KI-Zusammenfassung

GPT-6 Astra was benchmarked on Artificial Analysis Intelligence Index v4.3, which incorporates 10 evaluations including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The model's performance is tracked across capability indices for Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics, with metrics covering intelligence scores, cost per task, and output speed.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10