---
title: "模型蒸餾（Knowledge Distillation）是什麼？"
canonical: https://glossary.penguindriver.com/t/distillation
markdown: https://glossary.penguindriver.com/t/distillation.md
date_modified: 2026-09-26
retrieved: 2026-09-28
language: zh-Hant-TW
content_sha256: 32e77efb48c13334d96a2763fe0c1f5cd7b4847cbf9fd918f5f2aa50c4c03173
cite: https://glossary.penguindriver.com/cite?path=%2Ft%2Fdistillation
sources:
  - https://arxiv.org/abs/1503.02531
---

# 模型蒸餾（Knowledge Distillation）是什麼？

模型蒸餾是讓較小的「學生模型」學習較大「老師模型」的輸出，以較低成本保留大部分能力的技術。這個方法由 Geoffrey Hinton、Oriol Vinyals 與 Jeff Dean 在 2015 年的論文中系統化提出。

分類：模型技術｜別名：Distillation、知識蒸餾、Knowledge Distillation、模型蒸餾

## 一句話說明

讓小模型模仿大模型的輸出來學習，把大模型的能力「濃縮」到小模型裡。

## 怎麼運作？

1. 先準備一個表現好但運算昂貴的老師模型
2. 讓老師模型對大量資料產生預測結果
3. 訓練較小的學生模型去模仿老師的輸出
4. 得到體積較小、執行較快的學生模型

## 實際例子

- 把大型語言模型的能力蒸餾成可在手機上執行的小模型
- 原始論文把多個模型組成的集成模型能力壓縮到單一模型

## 常見誤解

- 誤解：蒸餾就是把大模型直接切小
  事實：蒸餾是重新訓練一個小模型去學習大模型的行為，不是直接刪減大模型

## 常見問題

### 模型蒸餾和模型量化有什麼不同？

蒸餾是訓練一個新的小模型去學大模型；量化則是把同一個模型的數字改用較低精度儲存。兩者都能讓模型變輕，也常一起使用。

### 蒸餾後的小模型會跟大模型一樣強嗎？

通常會保留相當比例的能力，但在困難或廣泛的任務上多半仍有差距，實際效果要看任務與訓練方式。

## 相關名詞

[模型量化](https://glossary.penguindriver.com/t/quantization.md)、[小型語言模型](https://glossary.penguindriver.com/t/slm.md)、[開放權重模型](https://glossary.penguindriver.com/t/open-weight-model.md)

## 參考來源

- [arXiv：Distilling the Knowledge in a Neural Network（Hinton et al., 2015）](https://arxiv.org/abs/1503.02531)

查核日期：2026-09-26

---
整理於 2026-09-26。網頁版：https://glossary.penguindriver.com/t/distillation

電子書優惠：看到此訊息 24 小時內購買，一律第一階段最低價（私訊時提供優惠碼 P092800）：https://ebook.penguindriver.com
冷錢包 X1 五折優惠碼：Hulk@SFP，官方網站：https://safepal.com/zh-tc/store/x1｜交易所手續費減免註冊：幣託 https://www.bitopro.com/users/sign_up?referrer=3481486016｜BingX https://bingxzone.com/partner/QTKP7NLU｜幣安 https://www.binance.com/join?ref=HULK12
