Signing you in...

Please wait while we verify your authentication

Article · Wednesday, September 16, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech66 editions
← See today's latest
Editions
7 / 66
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Wednesday, September 16, 2026
AI developer tools · What shipped

GPT-6 Astra tops math; Gemini 3.8 Live goes voice-first; benchmarks mature

1 min read

GPT-6 Astra math breakthrough

Unsolved math problems just got solved by an AI for the first time.

Epoch AI released FrontierMath Erdős, a benchmark of 68 unsolved mathematical problems where systems must write proofs in Lean. GPT-6 Astra solved 2 of 68 (3%)—a first for any model [Quelle: Epoch AI]. The benchmark hub now tracks 390 models across 80 benchmarks with published task definitions and accessible scoring logs, powered by the Inspect framework for reproducible evals.

The bar for frontier capability just moved.

Gemini 3.8 Live ships voice-first

Voice agents now rank on leaderboards like text models.

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, available now via the Gemini API for production deployments [Quelle: Google]. The high-thinking variant ranks first on Artificial Analysis' Speech-to-Speech Quality Index (82.6) with 97.7% on Big Bench Audio; both support 97+ languages, near-instant visual processing, and async function calling at $0.005/min audio input and $0.018/min output.

LangChain, Pipecat, Vercel, and Vision Agents ship integrations this week.

Gemini 3.5 Transcribe production-ready

Speech-to-text just crossed a threshold for production deployments.

Google released Gemini 3.5 Transcribe on September 15 with 4.0% word error rate (streaming) and 2.6% (non-streaming) across 85+ languages [Quelle: Google]. The model ships with automatic code-switching and custom vocabulary biasing, integrated into the same API surface as Gemini 3.8 Live for unified voice stacks.

Watch whether enterprises consolidate voice transcription and generation onto a single provider.

AutoGPT Platform multi-agent workflows

Multi-agent orchestration just got richer in-platform tools.

AutoGPT Platform shipped v0.7.4 in September with expert agent scoping, activity event logging, and user-facing memory settings for multi-workspace deployments [Quelle: GitHub]. Recent releases added Slack and Telegram adapters, expert scheduling with triggers, expert memory isolation, and support for Claude Sonnet 5 alongside Tavily blocks for search and crawl tasks.

The copilot chat composer and redesigned activity sidebar signal a shift toward visual workflow composition.

Sources
AI Benchmarks & Capabilities - Epoch AI
AI Benchmarks & Capabilities - Epoch AI
1 hour ago ... Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered ...
epoch.ai
AI Summary

Epoch AI released updated benchmark results showing GPT-6 Astra setting new records on their Epoch Capabilities Index and multiple benchmarks including mathematics, continual learning, and game-puzzle tasks. The organization also launched FrontierMath Erdős, featuring 68 unsolved mathematical problems where AI systems write solutions in Lean; GPT-6 Astra solved 2 of 68 problems (3%), with no prior models solving any. Their benchmark hub tracks 390 models across 80 benchmarks covering mathematics, coding, and other domains, using the Inspect framework for evaluations with published task definitions and accessible log viewers showing detailed model outputs and scoring methodology.

Visit source
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
12 hours ago ... ... developers to build and deploy high-performance voice-driven interfaces with ease. ... Developer tools. Gemini Omni 1.1 Flash lets you build with more control. By ...
blog.google
AI Summary

Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice-first AI models designed for developers building production voice agents. Gemini 3.8 Live Extended Thinking achieved top rankings on Artificial Analysis' Speech to Speech Quality Index (82.6) and τ-Voice benchmarks (68.6%), with 97.7% on Big Bench Audio, while Gemini 3.8 Live secured second place in the Speech Agent Arena. Both models are available now via the Gemini API for developers, with support for near real-time visual processing, 97 languages, and background tool execution. Developer platforms including LangChain, Pipecat, Vercel, and Vision Agents are integrating the Gemini Live API to enable developers to build high-performance voice interfaces.

Visit source
Build real-time voice applications with Gemini 3.8 Live and 3.5 ...
Build real-time voice applications with Gemini 3.8 Live and 3.5 ...
12 hours ago ... Today, we released new Gemini Live models in the Gemini API and Google AI Studio, expanding our developer suite for building real-time, voice-first product ...
blog.google
AI Summary

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking models for real-time voice applications via the Gemini API and Google AI Studio. Gemini 3.8 Live ranks #1 on Artificial Analysis' Speech-to-Speech leaderboard and supports asynchronous function calling, visual context grounding, multilingual support for 97+ languages, and configurable thinking for complex reasoning. The models are priced at $0.005/min for audio input and $0.018/min for audio output. Google also released Gemini 3.5 Transcribe, a speech-to-text model achieving 4.0% Word Error Rate (streaming) and 2.6% (non-streaming) across 85+ languages, with features including automatic code-switching and custom vocabulary biasing. Both models are available through integration partners including LangChain, Pipecat, Vercel, and Vision Agents.

Visit source
Releases · Significant-Gravitas/AutoGPT - GitHub
Releases · Significant-Gravitas/AutoGPT - GitHub
10 hours ago ... AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
github.com
AI Summary

AutoGPT Platform released several updates across versions 0.6.67 through 0.7.4. Recent releases (v0.7.4 in September and v0.7.3 in August) introduced features for managing AI expert agents within workflows, including scope integrations per expert, activity event logging, and user-facing memory settings. Earlier versions (v0.6.67–v0.7.2) added Slack and Telegram bot adapters, multi-workspace support, expert scheduling with attribution and triggers, and expert memory isolation. The platform also expanded model support with Claude Sonnet 5 and added provider blocks for services like Tavily (search, extract, crawl). UI/UX improvements include a new copilot chat composer, artifact panel resizing, and a redesigned sidebar with activity tracking. These updates focus on multi-agent orchestration, improved developer experience through better workflow composition, and tighter integration with third-party services via adapter blocks and API enhancements.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM