AI developer tools · What shipped
For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.
Claude Fable 5.1 leads agent benchmarks; GPT-6 Astra solves unsolved math; Gemini 3.1 Pro launches
2 Min. Lesezeit
Claude Fable 5.1 agentic dominance
Fable 5.1 now runs 45% cheaper on agent workloads.
Anthropic's refresh delivers Terminal-Bench 4.0 at 55.8% (up from 42.0% in Fable 5) and CursorBench 3.2.0 at 73.4%, while slashing costs through 75% cheaper cache reads at $0.25 per million tokens [Quelle: Anthropic]. Early partners including Jane Street Capital and Cognition report better root-cause analysis for debugging and improved long-running problem-solving. The cost floor for production IDE agents just reset.
Claude Mythos 5.1 ships alongside with identical internals and 60% fewer false-positive safety flags.
GPT-6 Astra cracks unsolved math
GPT-6 Astra solved 2 of 68 unsolved Erdős problems.
Epoch AI's FrontierMath Erdős benchmark—where systems write Lean proofs for genuinely open mathematical questions—saw GPT-6 Astra reach 3% solve rate, marking the first model to solve any [Quelle: Epoch AI]. The benchmarking hub now tracks 391 models across 85 tasks with reproducible evaluation code through the Inspect framework and public log viewers showing individual model outputs. Frontier claims just became auditable.
Expect other labs to publish their own Erdős solve rates this week.
Gemini 3.1 Pro previews complex reasoning
Google shipped Gemini 3.1 Pro with upgraded problem-solving for complex tasks.
Available now in preview on Vertex AI and via the Gemini API, the model delivers noticeably better performance for multi-step reasoning and is paired with Gemini 3.1 Flash-Lite, a faster variant tuned for high-volume workloads like translation and content moderation [Quelle: Google Cloud]. Google also expanded agent tooling: Claude Opus 5 and Claude Sonnet 5 now run on the Agent Platform, xAI's Grok 4.6 shipped in Model Garden, Cloud Run added sandboxes for safely executing AI-generated code, and Apigee's Model Context Protocol (MCP) went generally available for API-to-agent transformations. The agent platform just got richer vendor coverage.
Watch whether enterprises consolidate agents across multiple model providers on a single platform.
GPT-5.1 lands in Copilot Studio
Microsoft made GPT-5.1 available in Copilot Studio for early testing.
Available now as an experimental model for U.S.-based customers in Power Platform early-access environments, GPT-5.1 brings improved adaptability in thinking time for both chat and reasoning modes alongside OpenAI's public release [Quelle: Microsoft]. The rollout is tagged for non-production testing to evaluate performance against existing models before general availability. Enterprise copilot builders can now benchmark reasoning quality before committing to production deployment.
Scale organizations will likely A/B test GPT-5.1 against Fable 5.1 within the week.
AI Benchmarks & Capabilities - Epoch AI2 hours ago ... Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered ...epoch.ai
GPT-6 Astra set new records on Epoch AI's Capabilities Index and multiple benchmarks including mathematics, continual learning, and game-puzzle tasks following pre-release access from OpenAI in early September 2026. Epoch AI launched FrontierMath Erdős, featuring 68 unsolved Erdős problems where AI systems write solutions in Lean; GPT-6 Astra solved 2 of 68 problems (3%), marking the first model to solve any of these problems. The Epoch benchmarking hub tracks 391 models across 85 benchmarks covering mathematics, coding, and software engineering, with detailed evaluation code available through the Inspect framework and publicly accessible log viewers showing how models answered individual questions.
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic3 hours ago ... ... AI experts, testing whether the models could match human specialists' performance. Mythos 5.1's capabilities are greater than those of Mythos 5. However ...anthropic.com

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, advancing AI capabilities for code generation and agentic software development tasks. Fable 5.1 achieves significantly higher performance across multiple developer-focused benchmarks including Terminal-Bench 4.0 (55.8% vs 42.0% for Fable 5), CursorBench 3.2.0 (73.4% vs 70.5%), and Humanity's Last Exam (65.0% with tools vs 63.8%), while reducing costs by approximately 25% for typical workloads and up to 45% for highly agentic tasks through 75% cheaper cache read pricing. The model demonstrates improved long-running problem-solving capabilities, better root-cause analysis for debugging, and enhanced performance on agentic terminal coding tasks, with early-access partners like Jane Street Capital, Cognition (Devin), and Shopify reporting substantial improvements in code quality and developer efficiency over predecessor models.
Google Cloud latest news and announcements21 hours ago ... ... performance AI agents with complete control. Call to Action: Register for ... This update allows developers to transform APIs into AI-ready tools using ...cloud.google.com

Gemini 3.1 Pro, Google Cloud's latest model, is now available in preview and represents a significant capability upgrade for complex problem-solving tasks. The model delivers noticeably improved performance and is accessible through Vertex AI, Gemini Enterprise, the Gemini API in Google AI Studio, Android Studio, and via CLI. Gemini 3.1 Flash-Lite, a faster and more cost-efficient variant, has also rolled out in preview for high-volume developer workloads like translation and content moderation while maintaining quality for complex tasks like UI generation and instruction-following. Google Cloud introduced several AI developer tool improvements including Claude Opus 5 and Claude Sonnet 5 from Anthropic now available on the Agent Platform, and xAI's Grok 4.6 added to Model Garden for coding and agentic workflows. The platform expanded code generation capabilities through Database Migration Service's AI-assisted code conversion using Gemini for translating legacy stored procedures and custom functions to PostgreSQL. Cloud Run now offers sandboxes in public preview for safely executing AI-generated code in isolated environments. Apigee's Model Context Protocol (MCP) is now generally available, allowing developers to transform APIs into AI-ready tools using OpenAPI specifications for secure agentic access to enterprise data.