← senthil.uk

LLM Radar

A live feed of LLM infrastructure releases, papers, models and tools. Refreshed every 30 minutes. Built on Cloudflare Workers + KV.

Last updated How this works →

Latest stable releases

Inference
ProjectVersionReleased
vLLM 0.31.0
SGLang 0.5.21
llama.cpp b11558
Ollama 0.40.3
Languages
ProjectVersionReleased
Go 1.27.2 —
Rust 1.99.0
Perl 5.44.0
Infrastructure
ProjectVersionReleased
OpenTofu 1.13.1
Terragrunt 1.1.6
tflint 0.64.0
Alpine Linux 3.24.2
PostgreSQL 18.6
Redis 8.10.2

Reading

New models · OpenRouter Models
  1. StepFun: Step 5 Preview1M ctx
  2. Anthropic: Claude Haiku 5.51M ctx
  3. Google: Nano Banana 2.165.5K ctx
  4. Mistral: Mistral Large 41M ctx
  5. inclusionAI: Ling 3.1 Flash262.1K ctx
  6. Pareto 26.10 Preview1M ctx
  7. OpenAI: GPT-6.1 Sol Pro1.1M ctx
  8. OpenAI: GPT-6.1 Sol1.1M ctx
openrouter.ai →
Trending · Hugging Face Models
  1. google/embeddinggemma-2feature-extraction · ♥ 1.7K
  2. jialinyyzz/humanizertext-generation · ♥ 1K
  3. Cloudflare/clefimage-text-to-text · ♥ 2K
  4. abenzerps/Qwen-Image-2.1-Uncensored-GGUFtext-to-image · ♥ 4K
  5. Qwen/Qwen-Image-2.1-Turbotext-to-image · ♥ 521
  6. Aleph-Alpha/Kolibri-1text-generation · ♥ 883
  7. Lightricks/LTX-2.5image-to-video · ♥ 7.3K
  8. Venastine-Research/Xing4.0-29B-A4B-GGUFtext-generation · ♥ 689
huggingface.co →
arXiv · LLM inference Inference
  1. SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
  2. TokenRouter: Efficient Serving System for Token-Level LLM Routing
  3. Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read
  4. Compile the Table: Query-Calibrated Operator Compression for Tabular In-Context Learning
  5. RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
arxiv.org →
vLLM Blog Inference
  1. vLLM Support for NVIDIA Vera Rubin NVL72: 7.8x Throughput over GB200 NVL72
  2. DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
  3. Taking vLLM Apart: A Practical Guide to Disaggregated Serving
  4. Watermarking in vLLM
  5. Announcing vllm-metal: Concurrent Serving on Apple Silicon
vllm.ai →
NVIDIA Inference Inference
  1. Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
  2. Benchmarking LLM Inference at Scale with AIPerf
  3. How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
  4. How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
  5. Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
developer.nvidia.com →
PyTorch Blog Inference
  1. Session-Aware Agentic Inference with NVIDIA Dynamo
  2. Building Spyre as a Native PyTorch Device
  3. Modernizing Table Batched Embeddings with FBTriton
  4. Evolution of the PyTorch Media Processing Landscape
  5. PyTorch Hardware Enablement: Updates from the Accelerator Integration Working Group
pytorch.org →
Together AI Inference
  1. Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
  2. Together Link: open models in the harness you already use. Start with one command today.
  3. How to train your own Jev for $17
  4. Canary rollouts: upgrade models in production without downtime
  5. How a global fintech scaled coding agent traffic with Dedicated Model Inference
www.together.ai →
The Go Blog Go
  1. Arch-specific SIMD in Go
  2. Platform-independent SIMD in Go
  3. Size-Specialized Memory Allocation
  4. Goroutine Leak Profiles
  5. Generic Methods
go.dev →
Rust Blog Rust
  1. Demoting i686 Windows targets to std-only
  2. Announcing Rust 1.99.0
  3. Announcing a Maintainer in Residence: Scott Schafer for the Cargo team
  4. GitHub Actions leaking secrets when Miri output is cached
  5. Be alert: targeted attacks on prominent Rustaceans
blog.rust-lang.org →
This Week in Rust Rust
  1. This Week in Rust 672
  2. This Week in Rust 671
  3. This Week in Rust 670
  4. This Week in Rust 669
this-week-in-rust.org →
Perl.com Perl
  1. NVIDIA Donates $12,000 to The Perl and Raku Foundation
  2. Announcing the Perl Toolchain Summit 2027
  3. HeroDevs Donates $10,000 to The Perl and Raku Foundation
  4. Alpha-Omega Donates USD 250,000 for Perl and CPAN Security
  5. Three ways to write a table in Podlite
www.perl.com →
blogs.perl.org Perl
  1. ANNOUNCE: Perl.Wiki V 1.56, CPAN::MetaCurator V 1.33
  2. Peta::NN - Surface Detail
  3. Minion: The Silent Engine That Turns Perl Into a Modern Architecture Platform
  4. This week in PSC (238) | 2026-10-05
  5. Perl documentation (including POD) with Markdown and Docbook
blogs.perl.org →
Cloudflare Blog Cloud
  1. Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  2. Introducing on-demand CPU and memory profiling with flamegraphs for Workers and Durable Objects
  3. Deno is joining Cloudflare
  4. Bridging technical depth and usability: The story behind Radar’s redesign
  5. Building an evidence-grounded agentic security operations harness on Cloudflare
blog.cloudflare.com →
AWS What's New Cloud
  1. AWS Security Hub now exports findings to S3 in CSV or JSON format
  2. Amazon EC2 R8gd instances are now available in additional regions
  3. Amazon EC2 R8g instances now available in additional regions
  4. Amazon Bedrock now supports reasoning summaries for OpenAI models
  5. Anthropic Claude Sonnet 5.5 and Claude Opus 5.5 are now available on Kiro in AWS GovCloud (US)
aws.amazon.com →
Google Cloud Blog Cloud
  1. What’s new with Google Data Cloud
  2. Modernizing Unstructured Data Workflows: Alteryx Live Query meets Google Cloud BigQuery
  3. Welcome to Gemini at Work 2026: Introducing the Gemini agent
  4. Empowering SMBs to do more with Gemini
  5. Innovation in Ireland: How Irish brands scale with Gemini Enterprise
cloud.google.com →
OpenTofu IaC
  1. OpenTofu v1.13.0
  2. An Introduction to OpenTofu Symbol Libraries!
  3. A Vision for Built-in Linting
  4. OpenTofu v1.12.0
  5. OpenTofu 1.12.0-beta1 is now available
opentofu.org →
HashiCorp IaC
  1. Terraform 1.16 completes Actions lifecycles and brings imports into child modules
  2. Terraform provider for Google Cloud 8.0 now generally available
  3. Secure AI agents with HashiCorp Boundary
  4. Simplify compliance with the native pre-written policy experience in HCP Terraform
  5. HCP Vagrant deprecation: important dates and migration guidance
www.hashicorp.com →
Hugging Face Blog AI
  1. Impactful scheduling for GPU clusters
  2. The model that didn't exist, so you made it yourself
  3. Multimodal open d1 decision models for the edge
  4. Introducing Falcon ASR
  5. One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
huggingface.co →
Simon Willison AI
  1. Dwarf Fortress uses version control now
  2. Python 3.15.0 added to actions/python-versions
  3. Quoting The New York Times
  4. Deno is joining Cloudflare
  5. Quoting Matthew Green
simonwillison.net →
Hacker News General
  1. We're unlocking the biggest mysteries of the clitoris
  2. YouTuber builds 'Flock-like' camera to track police vehicles, gets police visit
  3. apsw: Another Python SQLite Wrapper
  4. How to Build Wealth as a Career Person
  5. How I reverse engineered a commercial spatial audio effect
news.ycombinator.com →
LWN.net General
  1. Stable kernels 7.2.10 and 6.18.56
  2. Python 3.15 released
  3. [$] Adding kernel control-flow-integrity checking to GCC
  4. [$] The state of systemd: 2026 edition
  5. Let's Encrypt moving to 64-day certificate lifetimes in 2027
lwn.net →

Platform status