Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Mittwoch, 15. Juli 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech22 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
12 / 22
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Mittwoch, 15. Juli 2026
AI developer tools · What shipped

Fable 5 holds index lead; Faros and Cohere release eval frameworks

1 Min. Lesezeit

Fable 5 index lead holds

Anthropic's margin stays firm at the top.

Claude Fable 5 holds 161 on the Epoch Capabilities Index, one point ahead of GPT-5.5 Pro, and maintains its lead across all seven newly tracked evaluations spanning agentic work, cybersecurity, algorithm engineering, forecasting, and physics [Quelle: Epoch AI]. The index itself grew this month to include 13 fresh benchmarks, signaling that evaluation surface matters more than point margin—vendors are competing on breadth now, not just score.

Watch which new domains move to production first.

Faros rubric-based code eval

Scoring patches without running tests just shipped.

Faros released an evaluation framework using rubric-based scoring to rate AI coding models on real engineering tasks, validated against SWE-bench data [Quelle: Faros]. Gold patches scored a mean of 0.80; repeated judge runs correlated at 0.96 with ±0.02 variation. The framework achieved AUC rankings of 0.75–0.86 in distinguishing successful from failed patches, preserving quality distinctions that binary pass-fail benchmarks miss.

This unlocks eval in environments where test coverage is incomplete.

Tiny Aya multilingual stack

Cohere's multilingual release now ships eight artifacts.

Expedition Tiny Aya released Tiny Aya (70+ languages, open weights), Tiny Aya Math Edition (multilingual Math Olympiad benchmark), Kids Companion (multilingual safety benchmark), Tiny Aya Vision (parameter-efficient multimodal), Tiny Facade (4-bit quantized variants), and open-source interpretability tools using sparse autoencoders [Quelle: Cohere]. The core research landed at COLM 2026; teams published papers, datasets, and won Hugging Face Build Small hackathon entries.

Local multilingual deployment just became reproducible and benchmarked.

Quellen
Data on AI Capabilities and Benchmarking - Epoch AI
Data on AI Capabilities and Benchmarking - Epoch AI
7 hours ago ... Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated ...
epoch.ai
KI-Zusammenfassung

Claude Fable 5 achieved a new high score of 161 on the Epoch Capabilities Index (ECI), surpassing GPT-5.5 Pro by 1 point and marking the first time Anthropic has led the ECI in over a year. Epoch AI also recently expanded its benchmarking hub by tracking 13 new evaluations, with 7 incorporated into the ECI, while adding nine external benchmarks spanning agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics.

Quelle öffnen
AI model routing: How we score code without running tests
AI model routing: How we score code without running tests
8 hours ago ... The AI code evaluation framework behind our open vs. frontier model test ... model generated the patch. How does a rubric score differ from test pass ...
faros.ai
KI-Zusammenfassung

Faros released an AI model routing evaluation framework that uses rubric-based scoring to assess AI coding model performance on real engineering tasks. The framework evaluates patches against weighted checklists of task-specific criteria rather than relying solely on test execution, enabling evaluation of work where original environments or complete test coverage may be unavailable. Validation against SWE-bench datasets showed gold patches scored a mean of 0.80, repeated judge runs had correlation of 0.96 with ±0.02 variation, and the framework achieved AUC rankings of 0.75-0.86 in distinguishing successful from failed patches. The framework demonstrated a Spearman correlation of +0.94 between model success rates and average rubric scores on their failed attempts, indicating it preserves meaningful quality distinctions that binary pass/fail benchmarks cannot capture.

Quelle öffnen
Tiny Aya Expedition Drives Multilingual Innovation - Cohere
Tiny Aya Expedition Drives Multilingual Innovation - Cohere
12 hours ago ... High-performance models for agentic, multimodal, multilingual AI · Transcribe ... Agentic coding model, built for practical software engineering. Advanced ...
cohere.com
KI-Zusammenfassung

Cohere released Tiny Aya, an open-weight multilingual model supporting 70+ languages designed for local deployment on resource-constrained devices. The Expedition Tiny Aya research program produced multiple projects with released benchmarks and datasets: Tiny Aya Math Edition created a public multilingual Math Olympiad benchmark for evaluating mathematical reasoning; Kids Companion developed a multilingual safety benchmark for child-focused AI systems across dozens of languages; Cross-Lingual Word-Sense Disambiguation introduced new resources for multilingual semantic evaluation; and Language-Aware Quantization explored compression techniques preserving performance across diverse languages. Additional releases include Tiny Aya Vision (parameter-efficient multimodal extension), Tiny Facade (4-bit quantized variants for on-device tool use), and open-source interpretability tools using sparse autoencoders for analyzing multilingual representations. The core Tiny Aya research was accepted to COLM 2026, with multiple teams publishing papers, benchmarks, datasets, and winning entries in the Hugging Face Build Small Hackathon.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10