Du wirst angemeldet...

Bitte warte, während wir deine Anmeldung überprüfen

Artikel · Samstag, 3. Oktober 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

Von Marius BongartsTech81 Ausgaben
← Zur aktuellen Ausgabe
Ausgaben
5 / 81
Über Nacht von KI aus öffentlichen Quellen erstellt, täglich aktualisiert.
AI developer tools · What shipped
Samstag, 3. Oktober 2026
AI developer tools · What shipped

Inference engines hit 90% speedups; NVIDIA releases Nemotron MoE family

1 Min. Lesezeit

MetaInfer agentic optimization

AI-built inference engines just beat hand-tuned baselines by 2.3×.

Baseten shipped MetaInfer, a knowledge-only LLM inference engine generator that uses Claude Code and Fable 5 to auto-tune serving kernels [Quelle: Baseten]. On a single NVIDIA B200 with Qwen-3.6-35B, the generated engine VibeQwen achieved 1,792 tokens/sec decode (up 90% over vLLM) and 12ms time-to-first-token (2.3× faster). At concurrency 32, throughput climbed 71% to 10,307 output tokens/sec. Replicating the approach on SAM 3.1 image segmentation yielded 50% gains, proving the knowledge base transfers across modalities.

Serving layer optimization just became algorithmic.

Prime Inference serverless

Production-grade open-model serving is now cross-datacenter resilient.

Prime Intellect launched Prime Inference, separating prefill and decode on NVIDIA Blackwell GPUs with NVFP4 KV compression (+50% cache capacity) and BLHNC block-major layouts (-47% NVLink overhead) [Quelle: Prime Intellect]. GLM-5.3 on OpenRouter sustained 101 tokens/sec per user across 66 concurrent sessions with 100% uptime. The stack bundles vLLM, NVIDIA Dynamo, Mooncake, and FlashInfer, plus a native sparse-MLA attention kernel and structural-tag support for near-zero-error agent tool calls.

Serverless and reserved capacity just became production-grade for frontier models.

NVIDIA Nemotron & Cosmos 3

NVIDIA shipped three Nemotron tiers and physical AI models in one sweep.

The Nemotron family spans Lightning (3B active, Mamba-2 MoE, 1M context), Super (12B active, 5× throughput uplift), and Ultra (55B active for agentic workloads), each with Multi-Token Prediction and speculative-decoding drafters across vLLM, SGLang, TRT-LLM, and llama.cpp [Quelle: Hugging Face]. Cosmos 3 variants (Super: 32B+32B, Nano: 8B+8B for RTX PRO, Edge: 4B+2B for robotics) cover video reasoning and generation. Release includes 200+ open datasets, specialized ASR and document-parsing models, multimodal embeddings, and RAG rerankers—all with reproducible training recipes.

Open-weight foundation models just shipped with production breadth.

Quellen
Agentic inference optimization: 50-90% faster engines - Baseten
Agentic inference optimization: 50-90% faster engines - Baseten
8 hours ago ... Baseten is the kind of place where we have Slack channels like #ai-papers-discuss and #model-performance-reading-group , and it was in one of these channels ...
baseten.co
KI-Zusammenfassung

MetaInfer, a knowledge-only LLM inference engine generator, achieved substantial performance gains over state-of-the-art solutions through AI-assisted optimization. In a benchmark on Qwen-3.6-35B-A3B with NVFP4 precision on a single NVIDIA B200, an AI-generated engine called VibeQwen outperformed vLLM 0.25.1 by up to 90% on single-stream decode speed (1,792 vs 943 tokens per second) and delivered 2.3x faster time-to-first-token (12ms vs 28ms). At concurrency level 32, VibeQwen achieved 71% higher throughput (10,307 vs 6,030 output tokens per second). A second experiment applied the same approach to SAM 3.1 image segmentation, achieving 50% throughput improvement (91 images/second on H100 vs Meta's baseline), demonstrating the reusability of the optimization knowledge base across different model architectures and modalities.

Quelle öffnen
Prime Inference: Fast, Reliable Serving for Frontier Open Models
Prime Inference: Fast, Reliable Serving for Frontier Open Models
6 hours ago ... This scale pushed us to optimize for sustained performance, quality, and reliability, rather than benchmark speed alone. Our first public deployment, GLM-5.3, ...
primeintellect.ai
KI-Zusammenfassung

Prime released Prime Inference, a production model serving platform supporting both serverless endpoints and reserved capacity for frontier open-source models. The system separates prefill and decode operations on NVIDIA Blackwell GPUs, implements NVFP4 KV compression to increase cache capacity by 50%, and uses block-major (BLHNC) KV layouts to reduce NVLink transfer overhead by 47%. GLM-5.3 deployed on OpenRouter achieves 101 tokens per second per user at 66 concurrent sessions with 100% uptime. The stack combines vLLM, NVIDIA Dynamo, Mooncake, and FlashInfer, with contributions including a native sparse-MLA attention kernel for FlashInfer and structural-tag support in Dynamo for reliable tool-call generation with near-zero error rates in production agent serving.

Quelle öffnen
nvidia - Hugging Face
nvidia - Hugging Face
4 hours ago ... ... released under permissive licenses with the training recipes and evaluation frameworks that produced them. ... Nemotron Reinforcement Learning Collection ...
huggingface.co
KI-Zusammenfassung

NVIDIA released the Nemotron family of foundation models with open weights and reproducible training recipes optimized for various deployment scenarios. Key releases include Nemotron 3.5 Lightning (30B total / 3B active parameters with Mamba-2 + Transformer MoE architecture and 1M-token context), Nemotron 3 Super (120B total / 12B active with up to 5x higher throughput than previous versions), and Nemotron 3 Ultra (550B total / 55B active for demanding agentic workloads). These models feature Multi-Token Prediction layers and speculative-decoding drafters (DSpark, DFlash, MTP) for faster inference, with support across vLLM, SGLang, TRT-LLM, Ollama, and llama.cpp frameworks. NVIDIA also released Cosmos 3, an omni-model for physical AI with variants: Cosmos 3 Super (32B reasoner + 32B generator), Cosmos 3 Nano (8B + 8B optimized for RTX PRO workstations), and Cosmos 3 Edge (4B reasoner + 2B generator for real-time robotic deployment). The platform includes specialized models for speech recognition (Parakeet ASR, Canary multilingual), document parsing (Nemotron Parse), multimodal embeddings, and reranking for RAG pipelines, alongside 200+ open datasets covering pretraining, alignment, math reasoning, code generation, and evaluation benchmarks.

Quelle öffnen
Über Nacht zusammengestellt von MorningMail.aiZugestellt um 05:10