-
Updated
Aug 30, 2026 - Python
#
guidellm
Here are 4 public repositories matching this topic...
python benchmarking machine-learning performance gpu optimization cuda inference optuna local-first llm vllm guidellm
Artifact-backed LLM serving performance lab for vLLM baselines, official metrics, GuideLLM checks, and SGLang/PD scaffolding
python performance-engineering modal prometheus artifact-evaluation llm llm-serving vllm llm-inference sglang llm-performance gpu-benchmarking guidellm inference-benchmarking serving-metrics
-
Updated
May 21, 2026 - Python
Quantization from first principles + real-world vLLM serving benchmarks — BF16 vs on-the-fly FP8/INT8, measured with guidellm for memory, throughput, and latency trade-offs.
-
Updated
Jul 31, 2026 - Jupyter Notebook
Results parser for guideLLM with OpenSearch indexing capabilities
-
Updated
Aug 20, 2026 - Python
Improve this page
Add a description, image, and links to the guidellm topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the guidellm topic, visit your repo's landing page and select "manage topics."