Signing you in...

Please wait while we verify your authentication

Article · Saturday, October 3, 2026

AI developer tools · What shipped

For a senior engineer who already reads HN. Real changes in AI developer tools today: releases with version numbers, papers with benchmarks, repos that crossed a threshold worth knowing. Skip hype threads, pre-announcement leaks, and recycled summaries. Always link primary sources.

By Marius BongartsTech81 editions
← See today's latest
Editions
5 / 81
Generated by AI overnight from public sources, refreshed daily.
AI developer tools · What shipped
Saturday, October 3, 2026
AI developer tools · What shipped

Inference engines hit 90% speedups; NVIDIA releases Nemotron MoE family

1 min read

MetaInfer agentic optimization

AI-built inference engines just beat hand-tuned baselines by 2.3×.

Baseten shipped MetaInfer, a knowledge-only LLM inference engine generator that uses Claude Code and Fable 5 to auto-tune serving kernels [Quelle: Baseten]. On a single NVIDIA B200 with Qwen-3.6-35B, the generated engine VibeQwen achieved 1,792 tokens/sec decode (up 90% over vLLM) and 12ms time-to-first-token (2.3× faster). At concurrency 32, throughput climbed 71% to 10,307 output tokens/sec. Replicating the approach on SAM 3.1 image segmentation yielded 50% gains, proving the knowledge base transfers across modalities.

Serving layer optimization just became algorithmic.

Prime Inference serverless

Production-grade open-model serving is now cross-datacenter resilient.

Prime Intellect launched Prime Inference, separating prefill and decode on NVIDIA Blackwell GPUs with NVFP4 KV compression (+50% cache capacity) and BLHNC block-major layouts (-47% NVLink overhead) [Quelle: Prime Intellect]. GLM-5.3 on OpenRouter sustained 101 tokens/sec per user across 66 concurrent sessions with 100% uptime. The stack bundles vLLM, NVIDIA Dynamo, Mooncake, and FlashInfer, plus a native sparse-MLA attention kernel and structural-tag support for near-zero-error agent tool calls.

Serverless and reserved capacity just became production-grade for frontier models.

NVIDIA Nemotron & Cosmos 3

NVIDIA shipped three Nemotron tiers and physical AI models in one sweep.

The Nemotron family spans Lightning (3B active, Mamba-2 MoE, 1M context), Super (12B active, 5× throughput uplift), and Ultra (55B active for agentic workloads), each with Multi-Token Prediction and speculative-decoding drafters across vLLM, SGLang, TRT-LLM, and llama.cpp [Quelle: Hugging Face]. Cosmos 3 variants (Super: 32B+32B, Nano: 8B+8B for RTX PRO, Edge: 4B+2B for robotics) cover video reasoning and generation. Release includes 200+ open datasets, specialized ASR and document-parsing models, multimodal embeddings, and RAG rerankers—all with reproducible training recipes.

Open-weight foundation models just shipped with production breadth.

Sources
Agentic inference optimization: 50-90% faster engines - Baseten
Agentic inference optimization: 50-90% faster engines - Baseten
8 hours ago ... Baseten is the kind of place where we have Slack channels like #ai-papers-discuss and #model-performance-reading-group , and it was in one of these channels ...
baseten.co
AI Summary

MetaInfer, a knowledge-only LLM inference engine generator, achieved substantial performance gains over state-of-the-art solutions through AI-assisted optimization. In a benchmark on Qwen-3.6-35B-A3B with NVFP4 precision on a single NVIDIA B200, an AI-generated engine called VibeQwen outperformed vLLM 0.25.1 by up to 90% on single-stream decode speed (1,792 vs 943 tokens per second) and delivered 2.3x faster time-to-first-token (12ms vs 28ms). At concurrency level 32, VibeQwen achieved 71% higher throughput (10,307 vs 6,030 output tokens per second). A second experiment applied the same approach to SAM 3.1 image segmentation, achieving 50% throughput improvement (91 images/second on H100 vs Meta's baseline), demonstrating the reusability of the optimization knowledge base across different model architectures and modalities.

Visit source
Prime Inference: Fast, Reliable Serving for Frontier Open Models
Prime Inference: Fast, Reliable Serving for Frontier Open Models
6 hours ago ... This scale pushed us to optimize for sustained performance, quality, and reliability, rather than benchmark speed alone. Our first public deployment, GLM-5.3, ...
primeintellect.ai
AI Summary

Prime released Prime Inference, a production model serving platform supporting both serverless endpoints and reserved capacity for frontier open-source models. The system separates prefill and decode operations on NVIDIA Blackwell GPUs, implements NVFP4 KV compression to increase cache capacity by 50%, and uses block-major (BLHNC) KV layouts to reduce NVLink transfer overhead by 47%. GLM-5.3 deployed on OpenRouter achieves 101 tokens per second per user at 66 concurrent sessions with 100% uptime. The stack combines vLLM, NVIDIA Dynamo, Mooncake, and FlashInfer, with contributions including a native sparse-MLA attention kernel for FlashInfer and structural-tag support in Dynamo for reliable tool-call generation with near-zero error rates in production agent serving.

Visit source
nvidia - Hugging Face
nvidia - Hugging Face
4 hours ago ... ... released under permissive licenses with the training recipes and evaluation frameworks that produced them. ... Nemotron Reinforcement Learning Collection ...
huggingface.co
AI Summary

NVIDIA released the Nemotron family of foundation models with open weights and reproducible training recipes optimized for various deployment scenarios. Key releases include Nemotron 3.5 Lightning (30B total / 3B active parameters with Mamba-2 + Transformer MoE architecture and 1M-token context), Nemotron 3 Super (120B total / 12B active with up to 5x higher throughput than previous versions), and Nemotron 3 Ultra (550B total / 55B active for demanding agentic workloads). These models feature Multi-Token Prediction layers and speculative-decoding drafters (DSpark, DFlash, MTP) for faster inference, with support across vLLM, SGLang, TRT-LLM, Ollama, and llama.cpp frameworks. NVIDIA also released Cosmos 3, an omni-model for physical AI with variants: Cosmos 3 Super (32B reasoner + 32B generator), Cosmos 3 Nano (8B + 8B optimized for RTX PRO workstations), and Cosmos 3 Edge (4B reasoner + 2B generator for real-time robotic deployment). The platform includes specialized models for speech recognition (Parakeet ASR, Canary multilingual), document parsing (Nemotron Parse), multimodal embeddings, and reranking for RAG pipelines, alongside 200+ open datasets covering pretraining, alignment, math reasoning, code generation, and evaluation benchmarks.

Visit source
Compiled overnight by MorningMail.aiDelivered at 05:10 AM