Update supported models

This commit is contained in:
Jiang-Jia-Jun
2025-06-30 08:16:03 +08:00
committed by GitHub
parent 16b0b51a5d
commit 39ed715b5e

View File

@@ -67,7 +67,7 @@ Learn how to use FastDeploy through our documentation:
| Model | Data Type | PD Disaggregation | Chunked Prefill | Prefix Caching | MTP | CUDA Graph | Maximum Context Length |
|:--- | :------- | :---------- | :-------- | :-------- | :----- | :----- | :----- |
|ERNIE-4.5-300B-A47B | BF16/WINT4/WINT8/W4A8C8/WINT2/FP8 | ✅WINT4/W4A8C8/Expert Parallelism)| ✅ | ✅|✅(WINT4)| WIP |128K |
|ERNIE-4.5-300B-A47B-Base| BF16/WINT4/WINT8 | ✅WINT4/Expert Parallelism)| ✅ | ✅|✅(WINT4)| | 128K |
|ERNIE-4.5-300B-A47B-Base| BF16/WINT4/WINT8 | ✅WINT4/Expert Parallelism)| ✅ | ✅|✅(WINT4)| WIP | 128K |
|ERNIE-4.5-VL-424B-A47B | BF16/WINT4/WINT8 | WIP | ✅ | WIP | ❌ | WIP |128K |
|ERNIE-4.5-VL-28B-A3B | BF16/WINT4/WINT8 | ❌ | ✅ | WIP | ❌ | WIP |128K |
|ERNIE-4.5-21B-A3B | BF16/WINT4/WINT8/FP8 | ❌ | ✅ | ✅ | WIP | ✅|128K |