Skip to content
#

model-quantization

Here are 77 public repositories matching this topic...

A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.

  • Updated Sep 11, 2026

AI Engineering: Annotated NBs to dive into Self-Attention, In-Context Learning, RAG, Knowledge-Graphs, Fine-Tuning, Model Optimization, and many more.

  • Updated Apr 2, 2025
  • Jupyter Notebook

H.E.R.A. (Healing Evaluation and Recognition Architecture) — an on-device edge-AI system for offline wound assessment using YOLOv8 segmentation, HSV tissue analysis, and PUSH-based decision logic.

  • Updated Aug 30, 2026
  • Dart

🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.

  • Updated Mar 27, 2025
  • Shell

The Ark Project: Selecting the perfect AI model to reboot civilization from a 64GB USB drive. Comprehensive analysis of open-source LLMs under extreme constraints, with final recommendation: Meta Llama 3.1 70B Instruct (Q6_K GGUF). Includes interactive tools, detailed comparisons, and complete implementation guide for offline deployment.

  • Updated Aug 7, 2025
  • HTML

Add this topic to your repo

To associate your repository with the model-quantization topic, visit your repo's landing page and select "manage topics."

Learn more