AgentDish directory

inference

Accepted listings with this tag.

Listing Category Score Trend Checked
#11 ↓ -3
gemma4.c

A pure C runtime for Gemma 4 E2B CPU inference, with export, benchmarking, and numerical validation tools included.

Developer Tools / Machine Learning / Inference Runtime 91 ↓ -3 2 days ago Details
#80 → 0
vLLM v0.28.0

Release page for vLLM v0.28.0, a high-throughput and memory-efficient inference and serving engine for LLMs. The snapshot shows release highlights, model support updates, breaking changes, and installable artifacts for PyPI, Docker, ROCm, CPU, and XPU.

Developer Tools / LLM Serving / Inference 89 → 0 46 hours ago Details
#102 → 0
AirLLM

AirLLM is an open-source inference library that aims to run very large language models on small GPUs by decomposing models layer-wise and streaming them during execution. The snapshot shows a detailed README with quickstart code, model examples, supported model notes, and recent update history.

Developer Tools / AI Inference 89 → 0 28 days ago Details
#158 ↓ -79
AutoRound

AutoRound is an open-source quantization toolkit for LLMs and VLMs, focused on high-accuracy low-bit inference across CPU, XPU, CUDA, and multiple deployment backends.

Developer Tools / AI Infrastructure 89 ↓ -79 118 days ago Details

A web calculator for estimating whether local LLMs fit on specific hardware and how fast they may run. It covers VRAM, token throughput, latency, power, and cost across multiple GPUs and devices.

Developer Tools / Code Assistant 88 ↓ -3 25 days ago Details
#300 ↓ -3
grunden.ai

A Swedish AI inference API offering GLM 5.1 on owned H200 hardware in Stockholm, with OpenAI-compatible endpoints, SEK pricing, and an emphasis on EU data residency.

Developer Tools / API 88 ↓ -3 111 days ago Details
#623 ↓ -3
microgpt-c

A pure C implementation for training and running a tiny GPT, with build instructions, sample output, and performance notes for macOS, Linux, and Windows.

Developer Tool / Machine Learning / AI Framework 85 ↓ -3 12 days ago Details

A technical primer that explains how LLM inference works, covering weights in GPU memory, prefill vs. decode, the KV cache, and batching.

Writing / Copywriting 84 ↓ -6 7 days ago Details
#791 ↓ -6
Frontier Roles

A job board focused on AI engineering roles, with first-party listings for RAG, agents, evals, inference, and MCP-related work. The page shows live filters, location and seniority breakdowns, and a nightly refreshed feed of open roles.

Jobs / AI Engineering Jobs 84 ↓ -6 30 days ago Details

An open-source handbook for production LLM serving and inference at scale, covering GPU fundamentals, KV cache, batching, quantization, speculative decoding, and engines like vLLM, SGLang, and TensorRT-LLM.

Developer Tools / AI Infrastructure 84 ↓ -6 86 days ago Details
#1155 ↓ -2
shaide

Self-hosted, Kubernetes-native platform for distributed multi-model LLM inference. It includes an OpenAI-compatible API, installer flow, and infrastructure-as-code deployment with support for air-gapped environments.

AI Infrastructure / LLM Serving 82 ↓ -2 just now Details
#1348 ↑ +2
Scalattice

Scalattice is an inference API and hosting platform that lets developers point the OpenAI SDK at its OpenAI-compatible endpoint with minimal code changes. The page includes Python and curl examples, model IDs, pricing references, and basic troubleshooting.

Developer Tools / API 81 ↑ +2 35 days ago Details

A pull request adding AMD GPU support to tiny-vLLM through ROCm/HIP while keeping the existing CUDA build path unchanged. The snapshot describes the compatibility header, CMake option, architecture selection, and validation on multiple AMD GPUs.

Developer Tool / LLM inference engine 79 ↑ +2 63 days ago Details

An Apple Silicon–optimized inference build of Bonsai 1.7B with custom Metal kernels, benchmark results, quick-start instructions, and a bundled OpenAI-compatible server.

Developer Tools / Code Assistant 79 ↓ -208 117 days ago Details

A GitHub repository for a custom llama.cpp backend that runs Qwen3-0.6B GGUF models on an Axera AX8850 NPU accelerator, with build/run notes, environment flags, architecture details, and test status.

Developer Tools / Code Assistant 78 ↑ +6 5 days ago Details
#1742 ↓ -1
THROTTLE

An interactive simulation about how inference providers ration access to stronger LLMs when demand spikes. Users can compare the default industry rule with their own policy and explore who gets the strong model.

AI Product / LLM Routing / Inference Management 74 ↓ -1 3 days ago Details