Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Dienstag, 22. September 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech81 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
16 / 81
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Dienstag, 22. September 2026
AI developer tools · What shipped

Fable 5.1 cost cuts widen lead; agentic benchmarks standardize

2 Min. Lesezeit

Fable 5.1 & Mythos 5.1

Agentic workloads just got 45% cheaper overnight.

Anthropic's refresh cuts cache reads by 75% to $0.25 per million tokens while Fable 5.1 holds Terminal-Bench 4.0 at 60.9% and CursorBench 3.2.0 at 73.4% on multi-file refactoring [Quelle: Anthropic]. Mythos 5.1—the same model with 60% fewer false-positive safety checks—reaches 10x higher binding affinities on protein design and delivers GPU kernel speedups up to 2.5x, making it the first usable for real biotech workflows. Enterprise Frontier Safeguards enable zero-data-retention deployments rolling out this fall across Claude Code, Vertex, and Bedrock.

Production IDE agents and research labs have a new cost floor.

NVIDIA Nemotron benchmarking

Agentic AI just got a two-layer evaluation standard.

NVIDIA's Nemotron 3.5 Lightning hits 86% accuracy on PinchBench while completing tasks 30% faster than Qwen3.6 35B, demonstrating production gains in tool-calling chains [Quelle: NVIDIA]. The benchmark framework measures both step-level process scoring and end-to-end outcomes across metrics like tool-call precision, argument accuracy, and cost per success, with executable environment verification prioritized over reference-based judging. Real-world domain-specific evaluations built from production tickets and APIs now inform enterprise deployment decisions.

Comparability just replaced opinion.

Anthropic's R&D transparency

Frontier labs just published how fast they're building AI.

Anthropic revealed that Claude leads 26% of its R&D work end-to-end and 90% has at least AI collaboration, with oversight tracking ~30,000 research agents at 0.002% blocking rate [Quelle: Anthropic]. Three frameworks measure compute allocation, process automation, and safety overhead—with 6% of R&D compute dedicated to safety work. The company proposes embedding independent evaluators to verify metrics as a standard other labs could adopt.

Transparency infrastructure just became competitive advantage.

BullshitBench nonsense rejection

Models that reject broken premises score highest.

BullshitBench measures whether AI confidently continues nonsensical assumptions or flags them clearly, testing 100 questions across 13 nonsense techniques in software, finance, legal, medical, and physics domains [Quelle: GitHub]. Updated September 10, the V2 suite covers 21,400 responses from models including Claude Fable 5.1, GPT-6 Astra, and DeepSeek V4.1 Flash, evaluated by a three-judge panel using frontier models. Responses fall into clear rejection, partial challenge, or acceptance—surfacing which vendors train for epistemic honesty.

Reasoning quality just got a new dimension.

Quellen
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
3 hours ago ... ... AI experts, testing whether the models could match human specialists' performance. Mythos 5.1's capabilities are greater than those of Mythos 5. However ...
anthropic.com
KI-Zusammenfassung

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, positioning them as the world's most advanced models for coding and knowledge work. Fable 5.1 shows significant performance improvements across multiple benchmarks: Terminal-Bench-Science 0.1 (52.6% vs Fable 5's 24.7%), Terminal-Bench 4.0 agentic coding (60.9%), CursorBench 3.2.0 (73.4%), and Humanity's Last Exam (65.0% with tools). The model costs approximately 25% less than Fable 5 for typical workloads and up to 45% less for highly agentic work due to 75% reductions in cache read pricing ($0.25 per million tokens vs prior rates). Fable 5.1 demonstrates improved long-context reasoning, better root-cause analysis in debugging, and improved safeguard precision with 60% fewer false positives in cybersecurity tasks while now supporting defensive vulnerability discovery. Mythos 5.1, the same underlying model with modified safeguards for vetted researchers, shows advanced capabilities in molecular design (achieving 10x higher binding affinities than previous competition winners with 50% hit rates on protein design tasks), computational biology optimization (up to 2.5x speedup on GPU kernels), and scientific reasoning. Enterprise Frontier Safeguards enable zero-data-retention deployments on customer infrastructure, rolling out this fall across Claude Code, Enterprise, Platform, and major cloud providers.

Quelle öffnen
How to Evaluate AI Agents From Tool Calls to Task Completion
How to Evaluate AI Agents From Tool Calls to Task Completion
7 hours ago ... Why benchmarks are converging on tool use. The line between “calling a ... developers use the incredible suite of AI tools available at NVIDIA. Chris ...
developer.nvidia.com
KI-Zusammenfassung

NVIDIA Nemotron 3.5 Lightning achieved 86% accuracy on PinchBench while completing tasks 30% faster than Qwen3.6 35B at comparable accuracy, demonstrating performance gains in agentic AI task completion. The benchmark employs a two-layer evaluation framework measuring both step-level process scoring and end-to-end outcome scoring across tool-calling chains, with core metrics including task success rate, tool-call precision, argument accuracy, steps per success, and cost per success rolled up through a fixed hierarchy of benchmark, trial, task, turn, and step levels. Executable environment verification is prioritized over reference-based or LLM-as-judge approaches for benchmark comparability, with real-world domain-specific evaluations built from production tickets and APIs now standard in enterprise AI agent deployment decisions.

Quelle öffnen
Measurements for understanding the pace of AI development inside ...
Measurements for understanding the pace of AI development inside ...
2 hours ago ... In this post, we lay out measurement tools that can illuminate three critical aspects of AI development: ... released to secure AI's benefits while staying on the ...
anthropic.com
KI-Zusammenfassung

Anthropic published detailed measurements of AI development pace inside its research lab, revealing that as of August 2026, Claude "leads" 26% of Anthropic's AI R&D work (meaning it completes most tasks end-to-end from high-level prompts with human supervision), while over 90% of AI R&D work has at least AI collaboration involvement. The company introduced three measurement frameworks: an R&D Automation Index tracking how much AI builds subsequent AI models, an oversight system monitoring ~30,000 research agents with 100% action coverage and a 0.002% blocking rate, and compute allocation metrics showing 6% of AI R&D compute dedicated to safety work. Anthropic plans to embed independent third-party evaluators to verify these metrics regularly and proposes this transparency model as a standard other frontier AI developers could adopt.

Quelle öffnen
petergpt/bullshit-benchmark: BullshitBench measures ... - GitHub
petergpt/bullshit-benchmark: BullshitBench measures ... - GitHub
24 hours ago ... BullshitBench measures whether AI models challenge nonsensical prompts instead of confidently answering them, created by Peter Gostev.
github.com
KI-Zusammenfassung

This content is a GitHub repository page for "BullshitBench," a benchmark tool that measures whether AI models detect and reject nonsensical premises rather than confidently continuing with invalid assumptions. The benchmark was updated September 10, 2026, and includes results for models like DeepSeek V4.1 Flash, Claude Fable 5.1, and GPT-6 Astra, tested at low and maximum reasoning levels. The V2 suite covers 100 questions with 214 model/reasoning variants, totaling 21,400 responses, and evaluates responses across 13 nonsense techniques spanning software, finance, legal, medical and physics domains. Models are scored on whether they clearly reject broken premises, partially challenge them, or accept the nonsense, with evaluation conducted by a three-judge panel using Claude Sonnet 4.6, GPT-5.2 and Gemini 3.1 Pro Preview.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10