AI developer tools · What shipped
For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.
OpenAI cuts GPT-6 pricing 50%; Anthropic & Fireworks ship token-efficient models
2 Min. Lesezeit
OpenAI GPT-6 Sol & Luna
OpenAI just halved pricing on two new GPT-6 variants.
Sol and Luna launched today in the API with input costs at $2 and $0.10 per million tokens respectively—both down 50% from GPT-5.6 [Quelle: MarkTechPost]. On AutomationBench, Sol outperforms Claude Opus 5 at 11x lower cost per task; Luna matches Opus 5 and Fable 5 on DeepSWE coding while costing 93–96% less. Both ship with improved prompt caching—up to 90% discounts on cached reads and new developer dashboards for cache diagnostics.
The tier below Astra just reset cost-performance.
Claude Fable 5.1 & Mythos 5.1
Following yesterday's update, Fable 5.1 adds new safeguard depth.
The model now ships with enhanced cybersecurity defenses that reduce false positives by 60% and explicit support for vulnerability discovery in defensive work [Quelle: Anthropic]. Mythos 5.1—the identical model with permissive safeguards for vetted security and life sciences professionals—gains Enterprise Frontier Safeguards enabling zero-data-retention deployments on customer infrastructure rolling out this fall. Cache reads stay at $0.25 per million tokens, holding the 45% cost advantage for agentic workloads.
Security teams and biotech labs get production-grade models for the first time.
Fireworks Ember-1
Token-efficient reasoning just hit production at scale.
Fireworks Research released Ember-1, a model matching Kimi K3's quality while slashing token use by 40% through learned selective reasoning [Quelle: Fireworks]. Live A/B tests with production customers on coding tasks confirmed 35% fewer tokens per task at parity quality; one customer now runs Ember-1 in production. The model built on Fireworks Serverless Training and ships today as a research preview with enterprise training support to enable custom token-efficient variants.
Token reduction just moved from research into live inference.
Gemini 3.8 Flash & Muse Spark 1.3
Google and Muse stake claims on cost-performance efficiency.
Gemini 3.8 Flash lands three variants on Artificial Analysis' intelligence index with the high-tier model at 41 capability points, 291 tokens/second output speed, and competitive pricing via the medium tier at $0.93 per task [Quelle: Artificial Analysis]. Muse Spark 1.3 counters with two variants, peaking at 48 capability points on the same index and pricing the xhigh-speed tier at $1.37 per task. Both models evaluated on 10 benchmarks spanning finance, strategy, legal, healthcare, and engineering domains.
The middle tier just got denser with options.
Gemini 3.8 Flash: Release Intelligence, Performance & Price12 hours ago ... The Gemini 3.8 Flash release offers 3 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key ...artificialanalysis.ai

Artificial Analysis Intelligence Index v4.3.2 has been released, incorporating 10 evaluations including AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The benchmark tracks model performance across weighted cost per task, output speed, and capability indexes covering finance, strategy, legal, healthcare, engineering, and economics domains. Gemini 3.8 Flash is highlighted as positioned in the most attractive cost-performance quadrant.
OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing ...24 hours ago ... ... Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device. > /scale_b2b_pipelineSPONSOR. 1M+. Monthly readers. AI engineers, researchers and founders.marktechpost.com

OpenAI released GPT-6 Sol and GPT-6 Luna, two new models in the GPT-6 family positioned below the flagship Astra. Both models are live in the OpenAI API today. Sol and Luna API prices have been cut by 50% compared to GPT-5.6 pricing: Sol input/output costs $2/$10 per 1M tokens (down from $4/$20), while Luna costs $0.10/$0.50 per 1M tokens (down from $0.20/$1.20). On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2% at $0.27 per task, outperforming Claude Opus 5 at 11.1x lower cost. On DeepSWE v1.1 coding benchmarks, Sol scores 68.8% at roughly 80% lower cost than Claude Fable 5, while Luna scores 66.6% comparably to Opus 5 and Fable 5 at medium effort but costs 93-96% less per task. On OSWorld 2.0 offline, Sol matches Opus 5's performance at 80% lower cost. OpenAI improved prompt caching with up to 90% discounts on cached input reads and new developer controls including a caching dashboard, diagnostics tools, and explicit breakpoints.
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic3 hours ago ... ... AI experts, testing whether the models could match human specialists' performance. ... model releases. A small number of customers' custom integrations ...anthropic.com

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, with Fable 5.1 priced 25% lower than Fable 5 for typical workloads due to 75% reduction in cache read pricing ($0.25 per million tokens), and up to 45% savings for highly agentic work. Input pricing remains $10 per million tokens and output $50 per million tokens. Fable 5.1 demonstrates improved performance across multiple benchmarks including Terminal-Bench-Science 0.1 (52.6%), Terminal-Bench 4.0 (55.8%), CursorBench 3.2.0 (73.4%), and Humanity's Last Exam (60.9% with tools), outperforming Fable 5 across coding, knowledge work, and agentic tasks. The models ship with enhanced developer capabilities including improved safeguards that reduce false positives by 60% in cybersecurity contexts, support for vulnerability discovery (defensive work only), and new Enterprise Frontier Safeguards allowing zero-data-retention deployments on customer-controlled cloud infrastructure. Mythos 5.1 adds permissive safeguards for vetted cybersecurity and life sciences professionals through trusted access programs developed in partnership with the US government.
Introducing Ember-1 - Fireworks AI7 hours ago ... Your AI performance stack is Fireworks + Voyage AI · fireworks kimi logo lockup. Model Releases6/12/2026. Kimi K2.7 Code on Fireworks: Better Agents, Lower Cost ...fireworks.ai

Ember-1 is a new specialized model from Fireworks Research delivering Kimi K3's quality with 40% fewer tokens by learning to cut unnecessary reasoning while preserving essential thinking. The model was developed through 50+ training experiments across mathematics, coding, instruction following, conversation, search, tool use, and software engineering tasks, using Fireworks Serverless Training to move from research to launch faster. On external benchmarks including Doximity's Bedside Bench (500 clinical cases across 10 categories), Ember-1 set a new Pareto frontier for cost-per-task performance against GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5, achieving 35–50% token reduction compared to K3 at maximum reasoning effort across seven industry benchmarks including Terminal Bench, SWE-bench Verified, and DeepSWE 1.1. Live A/B tests with two production customers on coding workloads confirmed approximately 35% fewer tokens per task at comparable quality, with one customer now running Ember-1 in live production. Ember-1 is available today as a Research Preview release on Fireworks Serverless with a two-week access model, with training support also launching to enable enterprises to build customized token-efficient models.