AgentDish directory

model-compression

Accepted listings with this tag.

Listing Category Score Trend Checked
#60 → 0
AirLLM

AirLLM is an open-source inference library that aims to run very large language models on small GPUs by decomposing models layer-wise and streaming them during execution. The snapshot shows a detailed README with quickstart code, model examples, supported model notes, and recent update history.

Developer Tools / AI Inference 89 → 0 20 hours ago Details