Logo
Explore Help
Sign In
apps/FastDeploy
1
0
Fork 0
You've already forked FastDeploy
mirror of https://github.com/PaddlePaddle/FastDeploy.git synced 2025-12-24 13:28:13 +08:00
Code Issues Actions 2 Packages Projects Releases Wiki Activity
Files
17b414c2df4ed1f7e74f0177bfc307ea29c384b6
FastDeploy/custom_ops
History
co63oc b6edd15d55 fix scaled_gemm_f8_i4_f16_weight_quantize input (#3685)
2025-08-29 11:04:04 +08:00
..
cpu_ops
[Code Simplification] remove cum_offsets (#3410)
2025-08-18 20:21:25 +08:00
gpu_ops
fix scaled_gemm_f8_i4_f16_weight_quantize input (#3685)
2025-08-29 11:04:04 +08:00
iluvatar_ops
[Iluvatar GPU] Optimze attention and moe performance (#3234)
2025-08-08 10:51:24 +08:00
utils
support w4afp8 EP inference (#3044)
2025-08-25 11:27:45 +08:00
xpu_ops
[Feature][XPU] add custom kernels for mtp (#3537)
2025-08-25 10:14:17 +08:00
0001-DeepGEMM-95e81b3.patch
[feat] support fa3 backend for pd disaggregated (#2695)
2025-07-03 22:33:27 +08:00
MANIFEST.in
[LLM] First commit the llm deployment code
2025-06-09 19:20:15 +08:00
setup_ops_cpu.py
polish code with new pre-commit rule (#2923)
2025-07-19 23:19:27 +08:00
setup_ops.py
[Optimize]support machete weight only gemm (#3561)
2025-08-28 09:49:58 +08:00
Powered by Gitea Version: 1.25.2 Page: 1024ms Template: 6ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API