LLM Radar
A live feed of LLM infrastructure releases, papers, models and tools. Refreshed every 30 minutes. Built on Cloudflare Workers + KV.
Latest stable releases
Reading
New models · OpenRouter Models
- StepFun: Step 5 Preview
- Anthropic: Claude Haiku 5.5
- Google: Nano Banana 2.1
- Mistral: Mistral Large 4
- inclusionAI: Ling 3.1 Flash
- Pareto 26.10 Preview
- OpenAI: GPT-6.1 Sol Pro
- OpenAI: GPT-6.1 Sol
Trending · Hugging Face Models
- google/embeddinggemma-2
- jialinyyzz/humanizer
- Cloudflare/clef
- abenzerps/Qwen-Image-2.1-Uncensored-GGUF
- Qwen/Qwen-Image-2.1-Turbo
- Aleph-Alpha/Kolibri-1
- Lightricks/LTX-2.5
- Venastine-Research/Xing4.0-29B-A4B-GGUF
arXiv · LLM inference Inference
- SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
- TokenRouter: Efficient Serving System for Token-Level LLM Routing
- Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read
- Compile the Table: Query-Calibrated Operator Compression for Tabular In-Context Learning
- RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
vLLM Blog Inference
- vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72
- DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
- Taking vLLM Apart: A Practical Guide to Disaggregated Serving
- Watermarking in vLLM
- Announcing vllm-metal: Concurrent Serving on Apple Silicon
NVIDIA Inference Inference
- Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
- Benchmarking LLM Inference at Scale with AIPerf
- How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
- How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
- Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
PyTorch Blog Inference
- Session-Aware Agentic Inference with NVIDIA Dynamo
- Building Spyre as a Native PyTorch Device
- Modernizing Table Batched Embeddings with FBTriton
- Evolution of the PyTorch Media Processing Landscape
- PyTorch Hardware Enablement: Updates from the Accelerator Integration Working Group
Together AI Inference
- Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
- Together Link: open models in the harness you already use. Start with one command today.
- How to train your own Jev for $17
- Canary rollouts: upgrade models in production without downtime
- How a global fintech scaled coding agent traffic with Dedicated Model Inference
The Go Blog Go
- Arch-specific SIMD in Go
- Platform-independent SIMD in Go
- Size-Specialized Memory Allocation
- Goroutine Leak Profiles
- Generic Methods
Rust Blog Rust
- Demoting i686 Windows targets to std-only
- Announcing Rust 1.99.0
- Announcing a Maintainer in Residence: Scott Schafer for the Cargo team
- GitHub Actions leaking secrets when Miri output is cached
- Be alert: targeted attacks on prominent Rustaceans
This Week in Rust Rust
this-week-in-rust.org →Perl.com Perl
- NVIDIA Donates $12,000 to The Perl and Raku Foundation
- Announcing the Perl Toolchain Summit 2027
- HeroDevs Donates $10,000 to The Perl and Raku Foundation
- Alpha-Omega Donates USD 250,000 for Perl and CPAN Security
- Three ways to write a table in Podlite
blogs.perl.org Perl
- ANNOUNCE: Perl.Wiki V 1.56, CPAN::MetaCurator V 1.33
- Peta::NN - Surface Detail
- Minion: The Silent Engine That Turns Perl Into a Modern Architecture Platform
- This week in PSC (238) | 2026-10-05
- Perl documentation (including POD) with Markdown and Docbook
Cloudflare Blog Cloud
- Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
- Introducing on-demand CPU and memory profiling with flamegraphs for Workers and Durable Objects
- Deno is joining Cloudflare
- Bridging technical depth and usability: The story behind Radar’s redesign
- Building an evidence-grounded agentic security operations harness on Cloudflare
AWS What's New Cloud
- AWS Security Hub now exports findings to S3 in CSV or JSON format
- Amazon EC2 R8gd instances are now available in additional regions
- Amazon EC2 R8g instances now available in additional regions
- Amazon Bedrock now supports reasoning summaries for OpenAI models
- Anthropic Claude Sonnet 5.5 and Claude Opus 5.5 are now available on Kiro in AWS GovCloud (US)
Google Cloud Blog Cloud
- What’s new with Google Data Cloud
- Modernizing Unstructured Data Workflows: Alteryx Live Query meets Google Cloud BigQuery
- Welcome to Gemini at Work 2026: Introducing the Gemini agent
- Empowering SMBs to do more with Gemini
- Innovation in Ireland: How Irish brands scale with Gemini Enterprise
OpenTofu IaC
- OpenTofu v1.13.0
- An Introduction to OpenTofu Symbol Libraries!
- A Vision for Built-in Linting
- OpenTofu v1.12.0
- OpenTofu 1.12.0-beta1 is now available
HashiCorp IaC
- Terraform 1.16 completes Actions lifecycles and brings imports into child modules
- Terraform provider for Google Cloud 8.0 now generally available
- Secure AI agents with HashiCorp Boundary
- Simplify compliance with the native pre-written policy experience in HCP Terraform
- HCP Vagrant deprecation: important dates and migration guidance
Hugging Face Blog AI
- Impactful scheduling for GPU clusters
- The model that didn't exist, so you made it yourself
- Multimodal open d1 decision models for the edge
- Introducing Falcon ASR
- One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Simon Willison AI
- Dwarf Fortress uses version control now
- Python 3.15.0 added to actions/python-versions
- Quoting The New York Times
- Deno is joining Cloudflare
- Quoting Matthew Green
Hacker News General
- We're unlocking the biggest mysteries of the clitoris
- YouTuber builds 'Flock-like' camera to track police vehicles, gets police visit
- apsw: Another Python SQLite Wrapper
- How to Build Wealth as a Career Person
- How I reverse engineered a commercial spatial audio effect