Comparatif exhaustif des méthodes de quantification LLM#
Page centrale du guide quantification — Compare en un coup d'œil toutes les méthodes de quantification pour Large Language Models, avec tableaux détaillés, diagrammes décisionnels ASCII, graphiques comparatifs et matrice de compatibilité des outils.
Sommaire#
- Introduction
- Tableau comparatif principal
- Guides de choix
- Graphiques ASCII comparatifs
- Matrice de compatibilité outils
- Cross-références
- Références
Introduction#
Pourquoi comparer les méthodes ?#
La quantification des LLMs n'est pas un problème à solution unique. Le choix de la méthode dépend de quatre facteurs principaux :
| Facteur | Question clé | Impact |
|---|---|---|
| Cas d'usage | Fine-tuning ? Inférence locale ? Serving cloud ? | Détermine la famille de méthodes |
| Matériel cible | GPU datacenter ? CPU ? Mobile ? Edge ? | Limite les formats exploitables |
| Budget qualité | Lossless acceptable ? Perte tolérée ? | Détermine le nombre de bits |
| Contrainte mémoire | Combien de VRAM/RAM disponible ? | Définit le bitrate minimum |
Les grandes familles#
(entraînement)"] ROOT --> PTQ["PTQ
(post-training)"] ROOT --> FK["Format/Kernel
(inférence)"] QAT --> Q1["LLM-QAT"] QAT --> Q2["BitNet"] PTQ --> P1["GPTQ"] PTQ --> P2["QuIP#"] PTQ --> P3["HQQ"] PTQ --> P4["AWQ"] PTQ --> P5["AQLM"] PTQ --> P6["SpQR"] PTQ --> P7["SqueezeLLM"] PTQ --> P8["OmniQuant"] FK --> F1["GGUF
(llama.cpp)"] FK --> F2["Marlin"] FK --> F3["QServe"]
Comment choisir — résumé express#
| Tu veux... | Choisis... |
|---|---|
| Fine-tuner un grand modèle sur un petit GPU | QLoRA / NF4 |
| Inférence GPU rapide en INT4 | GPTQ + Marlin |
| Inférence CPU / edge / mobile | GGUF (llama.cpp) |
| Compression maximale (2-bit) | QuIP# ou AQLM |
| Serving cloud haute throughput | QoQ / QServe |
| Hardware FP8 natif (H100/Blackwell) | INT8/FP8 HW |
| Quantification sans calibration | HQQ (data-free) |
| Entraîner un modèle from scratch | BitNet b1.58 |
Tableau comparatif principal#
Légende : PPL Δ = variation de perplexité vs baseline FP16 (approximatif, varie selon le modèle). Vitesse = speedup d'inférence relatif vs FP16. Mémoire = réduction vs FP16. Facilité = 🟢 facile · 🟡 moyen · 🔴 complexe.
Les valeurs sont approximatives et dépendent du modèle, du hardware et de la configuration exacte.
QAT et fine-tuning#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| QAT (LLM-QAT) | QAT | 4-bit | 4-bit (KV4) | Data-free | +0.05–0.1 | ~2x | ~4x | 🟡 | PyTorch, Meta | 2305.17888 |
| bitsandbytes / NF4 (QLoRA) | QAT+PTQ | 4-bit (NF4) | FP16 | Non | ~0 | 1x (overhead) | ~4x | 🟢 | bitsandbytes, PEFT, HF | 2305.14314 |
| BitNet b1.58 | QAT (from scratch) | 1.58-bit (ternary) | INT8 | N/A (from scratch) | ~0 (matche FP16) | >10x (théorique) | >10x | 🔴 | Microsoft, TernaryLM | 2402.17764 |
PTQ — Méthodes INT8 / W8A8#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| PTQ générique | PTQ | INT8 | INT8 / FP16 | Oui (128–1024) | +0.01–0.5 | ~2x | ~2x | 🟢 | PyTorch, TensorRT | — |
| LLM.int8() | PTQ | INT8 | INT8 (mixed-precision) | Non | ~0 (zéro dégradation) | ~1.5x | 2x | 🟢 | bitsandbytes, HF Transformers | 2208.07339 |
| SmoothQuant | PTQ | INT8 | INT8 | Oui | ~0 | 1.56x | 2x | 🟢 | smoothquant, vLLM | 2211.10438 |
| Outlier Suppression+ | PTQ | INT4 / INT6 / INT8 | INT4 / INT6 / INT8 | Oui | ~0 (INT8) | ~1.5x | 2x | 🟡 | OS+ repo | 2304.09145 |
PTQ — Méthodes INT4 / weight-only#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| GPTQ | PTQ | INT3 / INT4 | FP16 | Oui | +0.01–0.1 (INT4) | 3.25–4.5x | ~4x | 🟡 | AutoGPTQ, vLLM, ExLlamaV2 | 2210.17323 |
| AWQ | PTQ | INT3 / INT4 | FP16 | Oui | +0.01–0.05 (INT4) | 3x | ~4x | 🟡 | AutoAWQ, vLLM, TinyChat | 2306.00978 |
| HQQ | PTQ | 1–8 bit | FP16 | Non (data-free) | +0.1–0.3 (3-bit) | ~1x | 2–16x | 🟢 | hqq, HF Transformers, vLLM | — (blog) |
| OmniQuant | PTQ+ | W2–W6 | A4–A16 | Oui (optimisable) | +0.02–0.2 | ~1.5x | 2–8x | 🟡 | OmniQuant repo | 2308.13137 |
| SqueezeLLM | PTQ | 3–4 bit | FP16 | Oui | +0.05–0.2 (3-bit) | 2.3x | ~4x | 🔴 | SqueezeLLM repo | 2306.07629 |
| SpQR | PTQ | 3–4 bit | FP16 | Oui | < 1% ppl | 1.15x | ~4x | 🔴 | SpQR repo | 2306.03078 |
PTQ — Compression extrême (≤ 3 bit)#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| QuIP / QuIP# | PTQ | 2–4 bit | FP16 | Oui | +0.05–0.3 (2-bit) | ~1.5x | 4–8x | 🔴 | quip-sharp, HF (partiel) | 2307.13304 / 2402.04396 |
| AQLM | PTQ+ | 2–3 bit | FP16 | Oui (apprentissage codebooks) | +0.1–0.5 (2-bit) | ~1x (≥ FP16) | 5–8x | 🔴 | AQLM repo, HF Transformers | 2401.06118 |
Formats et kernels d'inférence#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| GGUF / GGML | Format | 2–8 bit | FP16 | Variable (K: non, I: oui) | Variable | 1–2x (CPU) | 2–8x | 🟢 | llama.cpp, Ollama, LM Studio | éval: 2601.14277 |
| Marlin | Kernel | INT4 | FP16 | N/A | N/A (préserve qualité) | 2.8–4x (batch) | N/A | 🟡 | vLLM, marlin repo | 2408.11743 |
| QoQ / QServe | PTQ+ | INT4 (W4) | INT8 (KV4) | Oui | +0.05–0.1 | 1.2–3.5x | ~3x | 🔴 | QServe repo | 2405.04532 |
| KV Cache (KIVI) | PTQ | FP16 | FP16 | Non (tuning-free) | ~0 (KV cache) | 2.35–3.47x | 2.6x | 🟢 | KIVI repo | 2402.02750 |
Hardware natif et configurations#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| INT8 / FP8 HW | PTQ / HW | INT8 / FP8 | INT8 / FP8 | Oui (statique) / Non (dynamique) | ~0 (FP8) | 2–4x | 2x | 🟢 | TensorRT, vLLM, cuBLAS | 2209.05433 |
| FP4 / FP8 | PTQ / HW | FP4 (E2M1) / FP8 | FP4 / FP8 | Variable | +0.01–0.1 | 2–8x (Blackwell) | 2–8x | 🟡 | Blackwell SDK, vLLM, MXFP | 2209.05433 |
| INT4 / INT3 / INT2 | PTQ | INT4 / INT3 / INT2 | FP16 | Variable | +0.01–0.5 | 2–4x | 4–8x | 🟡 | GPTQ, AWQ, HQQ, llama.cpp | — (régime) |
| W4A8 / W8A8 | PTQ | W4 / W8 | A8 | Oui | +0.01–0.1 | 2–3.5x | 2–4x | 🟡 | SmoothQuant, QServe, OmniQuant | — (configuration) |
Techniques émergentes#
| Méthode | Type | Bits (W) | Bits (A) | Calibration | PPL Δ | Vitesse | Mémoire | Facilité | Outils | arXiv |
|---|---|---|---|---|---|---|---|---|---|---|
| Sub-byte | Technique | 1–4 bit | Variable | Variable | Variable | Variable | 4–16x | 🟡 | llama.cpp, kernels custom | — |
| LLC | PTQ+ | Variable | Variable | Oui | Variable | Variable | Variable | 🔴 | Recherche | — |
| HFQ | PTQ | Variable | Variable | Non (hash-based) | Variable | Variable | Variable | 🟡 | Recherche | — |
Guides de choix#
Diagramme décisionnel principal#
= 1 GPU 48GB
= perf FP16"] Q1 -->|"Inférence"| INF Q1 -->|"Serving cloud"| SERV HW -->|"GPU"| GPU_INT4 HW -->|"CPU / Edge"| CPU_EDGE INF -->|"GPU INT4"| GPU_INT4["**GPTQ + Marlin** (vLLM)
ou AWQ + Marlin"] INF -->|"CPU/Edge"| CPU_EDGE["**GGUF** (llama.cpp)
Q4_K_M = best ratio"] SERV -->|"3× cheaper, batch 16-128"| QSQ["**QoQ / QServe**"] GPU_INT4 --> MAX["**COMPRESSION MAXIMALE ?**
(< 3 bits par poids)"] MAX --> QUIP["**QuIP#** (2-bit, Hadamard + E8)"] MAX --> AQLM["**AQLM** (2-bit, Multi-codebook)"] MAX --> HQQ_C["**HQQ** (data-free)"] MAX --> BITNET["**BitNet b1.58** (from scratch, ternaire)"]
Arbre de décision détaillé#
load_in_4bit=True + PEFT"] FT_GPU -->|"NON"| FT_SCRATCH{"Entraîner from scratch ?"} FT_SCRATCH -->|"OUI"| BITNET_FT["✅ BitNet b1.58 (ternaire)"] FT_SCRATCH -->|"NON"| SMALL["Utiliser un modèle plus petit"] FT -->|"NON"| INF{"Tu fais de l'INFÉRENCE LOCALE (1 utilisateur) ?"} INF -->|"OUI"| GPU_LOCAL{"Sur GPU ?"} GPU_LOCAL -->|"OUI"| INT4_OK{"INT4 acceptable ?"} INT4_OK -->|"OUI"| GPTQ_M["✅ GPTQ + Marlin (vLLM) ou AWQ"] INT4_OK -->|"NON"| LLMINT8["✅ LLM.int8() (bitsandbytes)"] GPU_LOCAL -->|"NON"| GGUF["✅ GGUF / llama.cpp
Q4_K_M = sweet spot
IQ2_XXS = compression max"] INF -->|"NON"| SERVING{"Tu fais du SERVING CLOUD (multi-user, batch) ?"} SERVING -->|"OUI"| H100{"GPU H100/Blackwell (FP8 natif) ?"} H100 -->|"OUI"| FP8_HW["✅ INT8/FP8 HW (TensorRT)
ou QoQ/QServe (W4A8KV4)"] H100 -->|"NON"| GPTQ_BATCH["✅ GPTQ + Marlin (INT4 batch)
ou QServe (L40S/A100)"] SERVING -->|"NON"| COMP_MAX{"Tu veux la COMPRESSION MAXIMALE ?"} COMP_MAX -->|"OUI"| BIT2_OK{"2-bit viable ?"} BIT2_OK -->|"OUI"| QUIP_C["✅ QuIP# (Hadamard + E8 lattice)
ou AQLM (multi-codebook)"] BIT2_OK -->|"NON"| INT4_STD["INT4 (GPTQ/AWQ/HQQ)"] COMP_MAX -->|"Data-free ?"| HQQ_C2["✅ HQQ (pas de calibration)"] COMP_MAX -->|"NON"| KV{"Tu veux quantifier le KV CACHE ?"} KV -->|"OUI"| KIVI["✅ KIVI (2-bit, tuning-free)"] KV -->|"NON"| STD["Méthode weight-only standard"]
Matrice de décision rapide#
| Méthode | Mémoire | Facilité | Vitesse | Compression | Qualité |
|---|---|---|---|---|---|
| QLoRA/NF4 | ★★★★ | ★★★★★ | ★★ | ★★★ | ★★★★★ |
| GPTQ+Marlin | ★★★★ | ★★★★ | ★★★★★ | ★★★ | ★★★★ |
| AWQ | ★★★★ | ★★★★ | ★★★★ | ★★★ | ★★★★★ |
| GGUF (Q4_K_M) | ★★★★ | ★★★★★ | ★★ | ★★★ | ★★★★ |
| QuIP# | ★★★★★ | ★★ | ★★★ | ★★★★★ | ★★★ |
| AQLM | ★★★★★ | ★★ | ★★★ | ★★★★★ | ★★★ |
| QoQ/QServe | ★★★★ | ★★ | ★★★★★ | ★★★ | ★★★★ |
| INT8/FP8 HW | ★★ | ★★★★ | ★★★★★ | ★★ | ★★★★★ |
| BitNet b1.58 | ★★★★★ | ★ | ★★★★★ | ★★★★★ | ★★★★ |
| HQQ | ★★★★ | ★★★★★ | ★★ | ★★★ | ★★★ |
| LLM.int8() | ★★ | ★★★★★ | ★★★ | ★★ | ★★★★★ |
★ = 1 (faible) · ★★★ = moyen · ★★★★★ = excellent
Graphiques ASCII comparatifs#
Perplexité vs méthode (à 4-bit, approximation)#
Plus la barre est courte, meilleure est la qualité. Baseline FP16 = 0.
Perplexité à 2-bit (compression extrême)#
Réduction mémoire vs méthode#
Plus la barre est longue, plus la compression est importante. Baseline FP16 = 1x.
Vitesse d'inférence relative vs méthode#
Speedup approximatif vs FP16 baseline. Dépend du batch size et du hardware.
Temps de quantification (coût de la mise en œuvre)#
Temps approximatif pour quantifier un modèle 7B sur A100.
BitNet b1.58 : entraînement from scratch (non représentable sur cette échelle)
Matrice de compatibilité outils#
Quel outil supporte quelle méthode ?#
| Méthode | HF Transformers | vLLM | TensorRT-LLM | ExLlamaV2 | llama.cpp | Ollama | TRITON |
|---|---|---|---|---|---|---|---|
| GPTQ | ✅ AutoGPTQ | ✅ | ✅ | ✅ | ❌ | ❌ | |
| AWQ | ✅ AutoAWQ | ✅ | ✅ | ❌ | ❌ | ❌ | ✅ |
| NF4/INT8 | ✅ natif | ✅ | ❌ | ❌ | ❌ | ❌ | |
| GGUF | ❌ | ❌ | ❌ | ❌ | ✅ natif | ✅ natif | |
| Marlin | via GPTQ | ✅ natif | ❌ | ❌ | ❌ | ❌ | ✅ |
| QoQ/QServe | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | |
| HQQ | ✅ | ✅ | ❌ | ❌ | ❌ | ❌ | |
| QuIP# | partiel | ❌ | ❌ | ❌ | ❌ | ❌ | |
| AQLM | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | |
| SpQR | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | |
| INT8 | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ | |
| FP8 | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ | |
| FP4 | ❌ | partiel | ✅ (Blackwell) | ❌ | ❌ | ❌ | |
| BitNet b1.58 | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
Tableau de compatibilité détaillé#
| Méthode | HF Transformers | vLLM | TensorRT-LLM | ExLlamaV2 | llama.cpp | Ollama |
|---|---|---|---|---|---|---|
| GPTQ | ✅ AutoGPTQ | ✅ | ✅ | ✅ | ❌ | ❌ |
| AWQ | ✅ AutoAWQ | ✅ | ✅ | ❌ | ❌ | ❌ |
| bitsandbytes (NF4/INT8) | ✅ natif | ✅ | ❌ | ❌ | ❌ | ❌ |
| GGUF | ❌ | ❌ | ❌ | ❌ | ✅ natif | ✅ natif |
| Marlin | via GPTQ | ✅ natif | ❌ | ❌ | ❌ | ❌ |
| QoQ/QServe | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| HQQ | ✅ | ✅ | ❌ | ❌ | ❌ | ❌ |
| QuIP# | partiel | ❌ | ❌ | ❌ | ❌ | ❌ |
| AQLM | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
| SpQR | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| INT8 | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ |
| FP8 | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ |
| FP4 | ❌ | partiel | ✅ (Blackwell) | ❌ | ❌ | ❌ |
| BitNet b1.58 | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
Compatibilité hardware#
| Hardware | Méthodes recommandées | Bits optimaux |
|---|---|---|
| NVIDIA H100 (Hopper) | FP8 natif, GPTQ+Marlin, QServe | FP8 / INT4 |
| NVIDIA A100 (Ampere) | GPTQ, AWQ, SmoothQuant, QServe | INT4 / INT8 |
| NVIDIA L40S | QoQ/QServe (optimisé), GPTQ+Marlin | INT4 / W4A8 |
| NVIDIA Blackwell (B200) | FP4 natif, FP8, INT4 | FP4 / FP8 |
| AMD MI300x | FP8 (ROCm), INT8 | FP8 / INT8 |
| GPU consumer (RTX 3090/4090) | GPTQ, AWQ, ExLlamaV2 | INT4 |
| CPU (x86/ARM) | GGUF / llama.cpp | Q4_K_M / IQ2–IQ4 |
| Apple Silicon (M-series) | GGUF / llama.cpp (Metal) | Q4_K_M / Q5_K_M |
| Mobile / Edge | GGUF, TFLite | Q4_K_S / INT8 |
Synthèse comparative par dimension#
Qualité (perplexité) par nombre de bits#
| Bits | Excellent (~0 PPL Δ) | Bon (< 0.1) | Acceptable (0.1-0.3) | Dégradé (> 0.3) |
|---|---|---|---|---|
| INT2 | QuIP# | AQLM, HQQ, GPTQ | ||
| INT3 | AWQ, GPTQ | HQQ | ||
| INT4 | AWQ, GPTQ | HQQ | ||
| INT8 | LLM.int8, SmoothQ, FP8 | |||
| FP16 | FP16 |
Trade-off qualité vs compression#
Cross-références#
Pages individuelles détaillées de chaque méthode dans le Brain KB.
QAT et fine-tuning#
PTQ — INT8 / W8A8#
PTQ — INT4 / weight-only#
Compression extrême#
Formats et kernels#
Hardware et configurations#
- INT8 / FP8 Hardware
- FP4 / FP8 Floating-Point
- INT4 / INT3 / INT2 — Quantification extrême
- W4A8 / W8A8 Mixed Precision
Techniques émergentes#
Références#
Papiers arXiv cités#
- LLM-QAT — Liu et al., LLM-QAT: Data-Free Quantization Aware Training for Large Language Models, arXiv:2305.17888, 2023
- PTQ générique — Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, arXiv:1712.05877, 2017
- LLM.int8() — Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, arXiv:2208.07339, 2022
- SmoothQuant — Xiao et al., SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, arXiv:2211.10438, 2022
- GPTQ — Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, arXiv:2210.17323, 2022
- AWQ — Lin et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, arXiv:2306.00978, 2023
- QLoRA — Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, arXiv:2305.14314, 2023
- QuIP — Chee et al., QuIP: 2-Bit Quantization of Large Language Models With Guarantees, arXiv:2307.13304, 2023
- QuIP# — Tseng et al., QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, arXiv:2402.04396, 2024
- AQLM — Egiazarian et al., Extreme Compression of Large Language Models via Additive Quantization, arXiv:2401.06118, 2024
- HQQ — Badri & Shaji, Half-Quadratic Quantization of Large Machine Learning Models, 2023 — blog technique
- SqueezeLLM — Kim et al., SqueezeLLM: Dense-and-Sparse Quantization, arXiv:2306.07629, 2023
- SpQR — Dettmers et al., SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression, arXiv:2306.03078, 2023
- OmniQuant — Shao et al., OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models, arXiv:2308.13137, 2023
- Outlier Suppression+ — Wei et al., Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling, arXiv:2304.09145, 2023
- MARLIN — Frantar et al., MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models, arXiv:2408.11743, 2024
- QServe / QoQ — Lin et al., QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving, arXiv:2405.04532, 2024
- KIVI — Liu et al., KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache, arXiv:2402.02750, 2024
- FP8 Formats — Micikevicius et al., FP8 Formats for Deep Learning, arXiv:2209.05433, 2022
- FP8 PTQ — Efficient Post-training Quantization with FP8 Formats, MLSys 2024
- BitNet — Wang et al., BitNet: Scaling 1-bit Transformers for Large Language Models, arXiv:2310.11453, 2023
- BitNet b1.58 — Ma et al., The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits, arXiv:2402.17764, 2024
- BitNet b1.58 2B4T — modèle open-source, arXiv:2504.12285, 2025
- GGUF evaluation — Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization Types, arXiv:2601.14277, 2026
- QQQ — QQQ: Quality Quattuor-Bit Quantization for Large Language Models, arXiv:2406.09904, 2024
- SpAtten — Wang et al., SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning, arXiv:2012.09852, 2020
- TernaryLM — implémentation alternative BitNet, arXiv:2602.07374, 2026
Sources complémentaires#
- Document de recherche source :
recherche-quantification-llm.md(Brain KB) - llama.cpp GitHub — documentation des formats GGUF
- Hugging Face Quantization Docs
- vLLM Quantization Docs
Glossaire rapide#
| Terme | Définition |
|---|---|
| WxAy | x bits pour les poids (Weight), y bits pour les activations (ex: W4A8) |
| KVx | KV cache en x bits |
| PTQ | Post-Training Quantization — quantification après entraînement |
| QAT | Quantization-Aware Training — entraînement avec quantification simulée |
| bpw | Bits per weight — bits par poids |
| GEMM | General Matrix Multiplication — multiplication matricielle |
| Outlier | Valeur d'activation avec magnitude exceptionnellement élevée |
| Calibration | Utilisation d'un petit dataset pour observer les distributions d'activation |
| Group-size | Taille du groupe de poids partageant un scale factor commun |
| Scale factor | Facteur d'échelle pour normaliser les valeurs avant quantification |
| Zero-point | Décalage pour la quantification asymétrique |
| STE | Straight-Through Estimator — propagation de gradients à travers la quantification |
| Incoherence | Propriété où poids et Hessienne ne sont pas alignés avec les axes coordonnés |
Page de navigation principale du guide quantification du Brain KB.
Dernière mise à jour : août 2026.