Signing you in...

Please wait while we verify your authentication

Article · Tuesday, July 14, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech22 editions
← See today's latest
Editions
13 / 22
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Tuesday, July 14, 2026
AI developer tools · What shipped

Open coding models arrive; Fable 5 holds Epoch lead

1 min read

Fable 5 benchmark lead

Claude Fable 5 extends its margin on Epoch's index.

Following previous issue, Fable 5 now scores 161 on the Epoch Capabilities Index, holding a one-point lead over GPT-5.5 Pro [Quelle: Epoch AI]. The index itself expanded this month to include seven fresh evaluations covering agentic work, cybersecurity, algorithm engineering, forecasting, and physics—domains where general benchmarks blind-spot. Fable 5 leads across all seven new evals.

The real signal is evaluation velocity, not point margin.

Open-weight coding models ship

Alibaba, Mistral, and Moonshot each released production coding models this month.

Alibaba's Qwen3-Coder and Qwen3-Coder-Next (80B MoE, 3B active per pass) target local agentic coding with lower cost per repo-level workflow [Quelle: Turing Post]. Mistral's Devstral 2 ships in dense (123B) and local-friendly (24B) flavors, both with 256K context and multi-file codebase manipulation. Moonshot's Kimi K2.7 Code is a 1-trillion parameter MoE with 256K context and autonomous multi-file execution.

Open weights are becoming competitive on speed and context window.

Proprietary coding models narrow

Grok 4.5 and Claude Sonnet 5 now cluster with Opus on production benchmarks.

Grok 4.5 ships with 80 TPS speeds and a 500k context window, benchmarking at Opus and GPT-5.5 tier [Quelle: Developers Digest]. Claude Sonnet 5 lands near Opus 4.8 on most tasks but at lower cost, though its new tokenizer runs approximately 30 percent more tokens per request. Pricing pressure from open weights is forcing velocity over margin.

Cost-per-correct-task replaces raw score as the real differentiator.

Sources
Data on AI Capabilities and Benchmarking - Epoch AI
Data on AI Capabilities and Benchmarking - Epoch AI
4 hours ago ... Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated ...
epoch.ai
AI Summary

Claude Fable 5 achieved a new high score of 161 on the Epoch Capabilities Index, surpassing GPT-5.5 Pro by 1 point and marking the first time Anthropic has led the index in over a year. Epoch AI recently expanded its benchmarking hub by tracking 13 new evaluations, with 7 incorporated into the Capabilities Index, and added nine external benchmarks spanning agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics.

Visit source
Best AI Coding Tools in 2026: Assistants, Agents, IDEs & Open Models
Best AI Coding Tools in 2026: Assistants, Agents, IDEs & Open Models
22 hours ago ... The category includes coding assistants, AI-native IDEs, terminal agents, repo-level agents, and open-source coding models that can run locally or inside ...
turingpost.com
AI Summary

Qwen3-Coder and Qwen3-Coder-Next have been released as open-weight coding models for agentic software work. Qwen3-Coder-Next is an 80B MoE model that activates 3B parameters per forward pass, designed for local coding agents and lower-cost repo-level workflows, with a research paper available. Kimi K2.7 Code is a 1-trillion parameter open-weight MoE model by Moonshot AI featuring 256K context window and autonomous multi-file execution capabilities. Devstral 2, released by Mistral AI, is an open-weight coding model family with dense (123B) and local-friendly (24B) variants, both with 256K context windows, purpose-built for multi-file codebase manipulation and tool use in autonomous software engineering agents.

Visit source
Developers Digest on AI Models
Developers Digest on AI Models
11 hours ago ... Meta's first paid API model arrives with $1.25/M input tokens, 1M context window, and strong tool-use benchmarks. HN debates what it means for the open-weights ...
developersdigest.tech
AI Summary

Grok 4.5 ships with 80 TPS speeds, a 500k context window, and benchmark results positioning it at Opus and GPT 5.5 tier, priced at $2/$6 per million tokens. Claude Sonnet 5 lands near Opus 4.8 on some tasks at a lower price point but runs approximately 30 percent more tokens due to a new tokenizer. GLM 5.2, an open-weight model, is positioned as a rival to GPT-5.5 with published benchmarks and pricing details.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM