Signing you in...

Please wait while we verify your authentication

Article · Monday, September 28, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech81 editions
← See today's latest
Editions
10 / 81
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Monday, September 28, 2026
AI developer tools · What shipped

Frontier models repriced; smaller open weights beat scale

1 min read

Claude Fable 5.1 & Mythos 5.1

Agentic coding just stayed the benchmark leader and got cheaper.

Following previous issue, Fable 5.1 now ships alongside Mythos 5.1, the variant with permissive safeguards for security and life sciences work [Quelle: Anthropic]. Mythos hits 50% on protein binder design tasks and runs GPU kernel speedups up to 2.5× for genomics pipelines; both models keep cache reads at $0.25 per million tokens (75% cheaper than the prior tier). Enterprise Frontier Safeguards unlock zero-data-retention on-prem deployments this fall.

Biotech teams finally have production-grade tools.

MiniCPM5-2B beats bigger rivals

A 2.5B-parameter open model just outscored larger competitors.

OpenBMB shipped MiniCPM5-2B in September, averaging 53.9 across 34 benchmarks—beating a 4B rival at 51.1—across coding, math, instruction following, and agentic tasks [Quelle: shattered.io]. The model supports a native 131k-token context window and deploys self-hosted with zero cloud costs. The same week saw Ornith-1.5-9B hit 70.6% on SWE-Bench coding-agent work and Liquid AI ship LFM2.5-VL-3B for speculative decoding.

Distillation and architecture efficiency just displaced raw parameter count.

GPT-6 Luna reprices frontier tier

OpenAI's newest tier cuts pricing across every workload axis.

GPT-6 Luna spans six models with intelligence scores ranging from 15 to 37 on Artificial Analysis Index v4.3.2, output speeds from 167 tokens/second, and pricing from $0.0045 per task at the low end [Quelle: Artificial Analysis]. Context window hits 1M tokens; 15× pricing spread lets teams pick effort level by workload. Benchmarks span Finance, Strategy, Legal, Healthcare, and Engineering domains.

Cost-performance just entered a new competitive era.

Benchmark disagreement widens between vendors

Published scores on identical tests now diverge by 20+ points.

Kimi K3 vs. Gemini 2.5 Pro show K3 ahead at 71.87 versus 50.24 overall, yet Gemini costs 40% less on chat turns ($0.00625 vs. $0.0105) and offers lower cache-heavy agent costs [Quelle: BenchLM]. Claude Sonnet 4.5 and Composer 2 show similar splits: Composer leads on Terminal-Bench agentic work at 61.7% while Sonnet costs 6× more per chat turn. Vendor benchmarks remain advisory only.

Golden-set evaluation runs are now table stakes before migration.

Sources
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
5 hours ago ... ... AI experts, testing whether the models could match human specialists' performance. ... model releases. A small number of customers' custom integrations ...
anthropic.com
AI Summary

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, setting new performance standards across multiple benchmarks. Fable 5.1 achieves 52.6% on Terminal-Bench-Science 0.1 (agentic scientific research), 60.9% on Terminal-Bench 4.0 (agentic coding), 60.9% on Humanity's Last Exam (multidisciplinary reasoning with tools), and 73.4% on CursorBench 3.2.0 (agentic coding)—outperforming Fable 5 and Opus 5 across these benchmarks. Pricing is reduced approximately 25% for typical workloads and up to 45% for highly agentic work due to 75% lower cache read costs ($0.25 per million tokens). The model demonstrates improved capabilities in coding, knowledge work, long-context reasoning, and scientific research, with Mythos 5.1 offering more permissive safeguards for cybersecurity and life sciences professionals. Fable 5.1 is available on all major platforms and the Claude API with production safeguards enabled.

Visit source
GPT-6 Luna Models - Intelligence, Performance & Price Comparison
GPT-6 Luna Models - Intelligence, Performance & Price Comparison
2 hours ago ... AI Trends · Premium · Log in. K. All Releases•. OpenAI logo. OpenAI. •. Proprietary release. •. Released September 2026. GPT-6 Luna: Release Intelligence, ...
artificialanalysis.ai
AI Summary

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, and others, measuring AI model performance across capability-specific indexes for Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics. The benchmark evaluates models on cost per task, output speed in tokens per second, and overall intelligence scores, with GPT-6 Luna positioned in the most attractive cost-performance quadrant.

Visit source
MiniCPM5-2B Beats Bigger Rivals by 2.8 Points [2026] - shattered.io
MiniCPM5-2B Beats Bigger Rivals by 2.8 Points [2026] - shattered.io
17 hours ago ... ... AI models like GPT-6 or Claude? Are these small-model benchmark claims independently verified? Related Coverage. A Crowded Week for Small AI Model Releases.
shattered.io
AI Summary

OpenBMB released MiniCPM5-2B in September 2026, a 2.5-billion-parameter language model that achieved an average score of 53.9 across 34 benchmarks covering coding, math, instruction following, general knowledge, long-context understanding, tool use, and agentic tasks, beating larger open-source rivals including a 4-billion-parameter model that scored 51.1. The model supports a native context window of 131,072 tokens and is designed for efficient self-hosted deployment. In the same week, Ornith AI released Ornith-1.5-9B, a 9-billion-parameter model achieving 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified for coding-agent tasks, while Liquid AI shipped LFM2.5-VL-3B-DSpark, a 279.5-million-parameter draft model for speculative decoding, and Supersonic Labs released Julia 1, a 144.3-million-parameter decision model. The releases reflect a broader shift toward squeezing capability from smaller parameter counts through distillation from frontier models, curated training data, and architecture optimizations focused on reasoning-per-parameter efficiency rather than raw scale. These developments are intensifying price competition with frontier models, as evidenced by DeepSeek's recent 70 percent output price cut and similar moves across the industry, while expanding the viability of on-device and edge deployment compared to larger hosted models.

Visit source
Gemini 2.5 Pro vs Kimi K3: Benchmarks & Cost | BenchLM.ai
Gemini 2.5 Pro vs Kimi K3: Benchmarks & Cost | BenchLM.ai
9 hours ago ... 3 results are shared. Category rows resting on Estimated evidence or different benchmark sets are marked directional and do not name a winner. Model A.
benchlm.ai
AI Summary

Kimi K3 and Gemini 2.5 Pro have been benchmarked across multiple AI evaluation frameworks with version numbers and public scores. Kimi K3 scores higher overall at 71.87 versus Gemini 2.5 Pro's 50.24, with particular strength in coding tasks (62.6 vs 24.6) and knowledge benchmarks like GPQA and HLE. Gemini 2.5 Pro offers lower API costs across standard workloads—chat turns at $0.00625 versus $0.0105, and cache-heavy agent loops at $0.15 versus $0.27—while Kimi K3 provides a larger documented context window at 1.05M tokens.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM