AI developer tools · What shipped
For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.
Claude Fable 5.1 deepens lead; Meta Muse, Ollama follow suit
1 Min. Lesezeit
Claude Fable 5.1
Anthropic's latest model tightens the cost-to-capability curve further.
Following yesterday's debut, Fable 5.1 now includes Claude Mythos 5.1, an unrestricted variant for verified cyberdefense and life sciences work [Source: Anthropic]. Protein binder design hit 50% success rates (typical baseline 10–15%), and computational biology runs 1.4–2.5× faster on deep learning workloads, cutting GPU costs by 30–60% on genome analyses. Safeguards now permit vulnerability discovery work while holding false positives to 60% lower rates in cybersecurity contexts.
Life sciences and security routing just got a legitimate frontier option.
Meta Muse Spark 1.3
Meta cuts 20% tool calls and 25% tokens, holding price steady.
Released two days ago, Muse Spark 1.3 achieves near-frontier performance on long-context retrieval (98.5 on MRCR) and coding benchmarks (75.4 on DeepSWE v1.1, 88.8 on Terminal-Bench 2.1) while maintaining $1.25/$4.25 per million token pricing [Source: Shattered]. A discounted contributor tier ($0.10/$0.20) trades training rights on your prompts for access. The million-token context window and efficiency gains make it a solid fit for cost-constrained teams handling long-horizon coding tasks.
Routing economics tighten again for agentic workloads.
Ollama v0.34 RC
Ollama ships ChatGPT Desktop integration and structured output gains.
Version 0.34-rc1 adds native ChatGPT Desktop support to run Ollama models directly in the app, alongside improved structured output performance on Apple Silicon [Source: GitHub]. Earlier v0.32.15 cut time-to-first-token from ~995ms to 524ms by caching resolved model metadata; v0.33.3 brought Gemma4 image and audio support on MLX. Recent additions include OpenAI-compatible client tool search and DeepSeek Harness integration for agentic coding tasks.
Local model latency and integration friction both keep dropping.
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic3 hours ago ... ... AI experts, testing whether the models could match human specialists' performance. Mythos 5.1's capabilities are greater than those of Mythos 5. However ...anthropic.com

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, advancing AI capabilities for coding and knowledge work. Fable 5.1 costs approximately 25% less than its predecessor for typical workloads (up to 45% less for agentic tasks) due to 75% reduced pricing on cache reads. The model demonstrates improved performance across multiple benchmarks: Terminal-Bench 4.0 (55.8%), Humanity's Last Exam with tools (65.0%), CursorBench 3.2.0 (73.4%), and OSWorld 2.0 (41.7% strict). Fable 5.1 shows stronger reasoning for debugging and root-cause analysis, with real-world validation from users including Jane Street Capital, Cognition, and others reporting superior performance over Fable 5 and Opus 5 across coding and long-running problem-solving tasks. Fable 5.1 introduces more precise safeguards reducing false positives by 60% in cybersecurity contexts while now permitting vulnerability discovery work. Claude Mythos 5.1, an unrestricted variant, is available through trusted access programs for vetted cyberdefenders and life sciences professionals. Scientific capabilities expanded to include protein binder design achieving 50% hit rates (versus typical 10-15%) and computational biology optimizations delivering 1.4x-2.5x speedups on deep learning models, reducing GPU costs by 30-60% on genome-wide analyses.
Releases · ollama/ollama - GitHub2 hours ago ... Enterprise platformAI-powered developer platform. AVAILABLE ADD-ONS. GitHub ... benchmarks); Fixes a bug where chat and generate could wedge after a mid ...github.com
Ollama released v0.34.0-rc1, adding ChatGPT Desktop integration to run Ollama models directly in the app, improving structured output performance on Apple Silicon, and adding support for OpenAI-compatible client tool search and response compaction. v0.33.3 brought Gemma4 image and audio support on MLX engine with MLX and llama.cpp updates. v0.32.15 improved time-to-first-token performance by caching resolved model metadata, reducing TTFT from approximately 995ms to 524ms in benchmarks. Recent releases also added support for Qwen3.8 27B with optimizations for Apple Silicon, DeepSeek Harness integration for agentic coding tasks, and structured output support across various model runners.
Meta Muse Spark 1.3 Cuts Tool Calls 20%, Tokens 25% [2026]19 hours ago ... Independent developers and benchmark sites will publish their own tool-call ... Nvidia Cuts AI Model Releases to 4-6 Weeks [2026] · Meta AI Glasses Cut ...shattered.io
![Meta Muse Spark 1.3 Cuts Tool Calls 20%, Tokens 25% [2026]](https://shattered.io/wp-content/uploads/2026/09/meta-muse-spark-1-3-agentic-coding-model-2026-1.webp)
Meta shipped Muse Spark 1.3 on September 2, 2026, an agentic coding model that achieves approximately 20% fewer tool calls and 25% fewer tokens than its predecessor Muse Spark 1.2 on equivalent tasks, according to Meta's internal engineering comparisons. The model supports a 1,048,576-token context window and maintains standard pricing at $1.25 per million input tokens and $4.25 per million output tokens, delivering an effective cost reduction despite the efficiency gains. On benchmarks, Muse Spark 1.3 scored 98.5 on MRCR long-context retrieval, 75.4 on DeepSWE v1.1, 88.8 on Terminal-Bench 2.1, and 1754 on GDPVal-AA v2 agent performance (compared to Opus's 1824), indicating optimization for coding and long-context workloads over general agent capability. A discounted contributor tier is available at $0.10 per million input tokens and $0.20 per million output tokens, requiring Meta training rights on prompts and completions.