FastDeploy

mirror of https://github.com/PaddlePaddle/FastDeploy.git synced 2025-12-24 13:28:13 +08:00

Author	SHA1	Message	Date
yzwu	ac013803f3	[Iluvatar] Support V1_KVCACHE_SCHEDULER and paddleocr-vl rope mode (#5555 )	2025-12-18 02:14:25 -08:00
Yuanle Liu	cdc0004894	Revert "[Feature] add ue8m0 for per_token_quant_fp8 (#5563 )" (#5611 ) This reverts commit `73e1d6aa90`.	2025-12-17 13:59:06 +08:00
Yuanle Liu	867803ae10	[BugFix] fix speculate_limit_thinking_content_length (#5590 ) * fix speculate_limit_thinking_content_length * update	2025-12-16 04:31:45 -08:00
chen	27ef3610b5	support glm fa3 (#5586 )	2025-12-16 19:33:27 +08:00
fxyfxy777	73e1d6aa90	[Feature] add ue8m0 for per_token_quant_fp8 (#5563 ) * ue8m0 * add default arg --------- Co-authored-by: YuBaoku <49938469+EmmonsCurse@users.noreply.github.com>	2025-12-16 18:40:12 +08:00
Echo-Nie	50100f98d7	[Feature] Support fusedmoe on Blackwell (#5325 ) * update sm100 * fix * fix style	2025-12-16 11:58:50 +08:00
freeliuzc	532f9ba227	[BugFix][Speculative Decoding](Spend many dyas to solve)Fix write qknorm cache bug in speculative decoding (#5491 ) * [liuzichang spend 10 dyas]fix write qknorm cache bug * fix 'fix cachekv bug''	2025-12-15 18:27:11 +08:00
chen	a389bb7c5c	[Feature][Optimization] Qwen Support Dynamic block_wise_fp8 cache (#5486 )	2025-12-12 17:10:17 +08:00
Juncai	d67388a479	[PD Disaggregation] Distinguish the pipelines for sending kv signal in different prefill (#5514 ) * Distinguish the pipelines for sending kv signal in different prefill * up	2025-12-12 14:05:36 +08:00
Neil Zhu	4403a21d4b	[Metax] refactor cutlass moe and optimize flash attention (#5361 ) * [Metax] refactor moe and flash attention backend --------- Co-authored-by: zhangchenyi_dl <16219492+zhangchenyidl@user.noreply.gitee.com>	2025-12-10 17:15:17 +08:00
Copilot	e38709b499	[BugFix] Fix limit_thinking early return logic in CUDA kernels (#5471 ) * Initial plan * [BugFix] Fix limit_thinking bug - change AND to OR in condition checks Co-authored-by: yuanlehome <23653004+yuanlehome@users.noreply.github.com> * Update Chinese comments to reflect OR logic instead of AND Co-authored-by: yuanlehome <23653004+yuanlehome@users.noreply.github.com> --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: yuanlehome <23653004+yuanlehome@users.noreply.github.com>	2025-12-10 11:03:19 +08:00
lzy	99f607eef5	[Others] Maintain the mtp branch temporarily. (#5446 )	2025-12-09 19:17:53 +08:00
lizexu123	95eab9f9ee	[Feature] support stop_token_ids (#5399 ) * support stop_token_ids * fix * delete chinese * support both * delete print	2025-12-09 17:49:12 +08:00
xiaozude	df67379bc3	[Metax] modify wrapSize to WARP_SIZE (#5442 )	2025-12-09 01:44:02 -08:00
周周周	31410415db	FA3 support qwen3 (#5441 )	2025-12-09 16:16:16 +08:00
K11OntheBoat	8d99bac532	Remove CUDA ERROR 9 of inputs of get_padding_offset kernel (#5440 ) Co-authored-by: K11OntheBoat <“ruianmaidanglao@163.com”>	2025-12-09 14:17:30 +08:00
周周周	2aea8a3a60	[Others] Remove useless code (#5404 )	2025-12-08 13:59:46 +08:00
GoldPancake	8545b705ed	fix top_p_candidates (#5400 ) Co-authored-by: freeliuzc <lzc842650834@gmail.com>	2025-12-05 20:01:05 +08:00
Yonghua Li	f4119d51b4	[PD Disaggregation] support DP via v1 router and decouple DP and EP (#5197 ) * [fix] support DP via v1 router and decouple DP and EP * [fix] fix scripts * [fix] reset model path * [fix] dp use get_output_ep, fix router port type, update scripts * [merge] merge with latest code * [chore] remove some debug log * [fix] fix code style check * [fix] fix test_multi_api_server for log_dir name * [chore] reduce logs * Apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>	2025-12-04 15:38:43 +08:00
周周周	a36d60aa18	[FIX BUG] fix bug in TP in permute_x_fp8_kernel (#5350 ) * commit * commit * commit * commit * commit * commit	2025-12-03 05:17:37 -08:00
Sunny-bot1	d5a9b75b4e	fix cutlass ep (#5337 )	2025-12-03 14:06:01 +08:00
lzy	c71a44c7e5	supports mtp split_kv_attn (#5343 )	2025-12-03 12:40:16 +08:00
Sunny-bot1	3629db4129	[Quantization] Support w4afp8 MoE dynamic quantization (#5282 ) * support dynamic activation quant for w4afp8 * support dynamic w4afp8 * add test * fix * fix --------- Co-authored-by: zhoutianzi666 <17801055074@163.com>	2025-12-02 18:56:16 +08:00
周周周	fb7f951612	[UNITEST] add test (#5305 )	2025-12-02 17:59:01 +08:00
K11OntheBoat	2e1680838f	[PD Disaggregation] Support PD deployment of DeepSeekv3. (#5251 ) * Support deepseekv3 cache transfer for PD deploy * clean some log info --------- Co-authored-by: K11OntheBoat <“ruianmaidanglao@163.com”>	2025-12-02 14:11:50 +08:00
chen	aa35ce449d	[Optimization] EP empty_input_forward Remove Communication (#5254 )	2025-12-01 21:10:40 +08:00
周周周	95243f012c	[Others] add PADDLE_ENFORCE (#5288 )	2025-11-28 14:23:35 +08:00
lizhenyun01	aba4fc657f	[Feature] support flash_mask_attention backend (#5134 ) * [Feature] suppert flash_mask_attention backend * fix unittest * clean code	2025-11-28 10:12:16 +08:00
GoldPancake	cfc5b0ccf9	[BugFix] fix mtp logprob bugs in chunk prefill (#5244 ) * fix mtp logprob bugs in chunk prefill * fix * fix	2025-11-27 11:31:29 +08:00
freeliuzc	ba915e03e1	[BugFix]Fix attention mask bug in D-Node of PD-split mode (#5245 )	2025-11-26 17:56:28 +08:00
xiaoxiaohehe001	61fc368066	[Fix] fix eplb noaux (#5239 ) * fix eplb noaux * fix eplb noaux	2025-11-26 17:50:51 +08:00
kevin	c068a4f642	[Feature] dyc8 support prefixcache (#5125 ) Some checks failed CE Compile Job / ce_job_pre_check (push) Has been cancelled Details CE Compile Job / print_ce_job_pre_check_outputs (push) Has been cancelled Details CE Compile Job / FD-Clone-Linux (push) Has been cancelled Details CE Compile Job / Show Code Archive Output (push) Has been cancelled Details CE Compile Job / BUILD_SM8090 (push) Has been cancelled Details CE Compile Job / BUILD_SM8689 (push) Has been cancelled Details CE Compile Job / CE_UPLOAD (push) Has been cancelled Details Deploy GitHub Pages / deploy (push) Has been cancelled Details Publish Job / publish_pre_check (push) Has been cancelled Details Publish Job / print_publish_pre_check_outputs (push) Has been cancelled Details Publish Job / FD-Clone-Linux (push) Has been cancelled Details Publish Job / Show Code Archive Output (push) Has been cancelled Details Publish Job / BUILD_SM8090 (push) Has been cancelled Details Publish Job / BUILD_SM8689 (push) Has been cancelled Details Publish Job / PADDLE_PYPI_UPLOAD_8090 (push) Has been cancelled Details Publish Job / PADDLE_PYPI_UPLOAD_8689 (push) Has been cancelled Details Publish Job / Run FD Image Build (push) Has been cancelled Details Publish Job / Run FastDeploy Unit Tests and Coverage (push) Has been cancelled Details Publish Job / Run FastDeploy LogProb Tests (push) Has been cancelled Details Publish Job / Extracted partial CE model tasks to run in CI. (push) Has been cancelled Details Publish Job / Run Base Tests (push) Has been cancelled Details Publish Job / Run Accuracy Tests (push) Has been cancelled Details Publish Job / Run Stable Tests (push) Has been cancelled Details CI Images Build / FD-Clone-Linux (push) Has been cancelled Details CI Images Build / Show Code Archive Output (push) Has been cancelled Details CI Images Build / CI Images Build (push) Has been cancelled Details CI Images Build / BUILD_SM8090 (push) Has been cancelled Details CI Images Build / Run FastDeploy Unit Tests and Coverage (push) Has been cancelled Details CI Images Build / Run FastDeploy LogProb Tests (push) Has been cancelled Details CI Images Build / Extracted partial CE model tasks to run in CI. (push) Has been cancelled Details CI Images Build / Run Base Tests (push) Has been cancelled Details CI Images Build / Publish Docker Images Pre Check (push) Has been cancelled Details * dyc8 support prefixcache * fix cache_trans test case * update code	2025-11-21 19:46:26 +08:00
freeliuzc	2d1dade5e2	[Speculative Decoding][MTP] Support static CacheKV C8 quantization and optimize memory usage (#5155 ) * support static cachekv c8 quantization in mtp mode * optimize memory allocation	2025-11-21 15:10:13 +08:00
xiaoxiaohehe001	6ca2651995	[Feature] Support noaux for eplb (#5143 ) * support noaux eplb * noaux_eplb * noaux_eplb * noaux_eplb	2025-11-21 14:10:32 +08:00
Yonghua Li	43097a512a	[BugFix] [PD Disaggregation] fix v1 scheduler prefill node profile run & ipc transfer protocol (#5132 ) Some checks failed CE Compile Job / ce_job_pre_check (push) Has been cancelled Details CE Compile Job / print_ce_job_pre_check_outputs (push) Has been cancelled Details CE Compile Job / FD-Clone-Linux (push) Has been cancelled Details CE Compile Job / Show Code Archive Output (push) Has been cancelled Details CE Compile Job / BUILD_SM8090 (push) Has been cancelled Details CE Compile Job / BUILD_SM8689 (push) Has been cancelled Details CE Compile Job / CE_UPLOAD (push) Has been cancelled Details Deploy GitHub Pages / deploy (push) Has been cancelled Details * [fix] fix v1 scheduler profile run for append attention in prefill node * [fix] skip send_signal if kv signal not inited for gpu and xpu * [fix] extend fix to flash_attn & mla_attn * [fix] fix v1 pd run in ipc transfer protocol * [ci] add test for v1 pd profile run using ipc transfer protocol * [style] fix code style check * [style] fix code style again * [fix] fix profile run * [update] remove --num-gpu-blocks-override in example script * [chore] rename forward_meta is_profiling to is_dummy_or_profile_run	2025-11-20 21:39:22 +08:00
Jundong Liu	147b2e5eb0	[BugFix] Fix zero workspace returned by CUB size query under CUDA Graph in MoE dispatch (#5087 ) * fix bug about CubKeyValueSorter::run * pre-commit and add comment * pre-commit * Apply suggestion from @Copilot Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * fix precommit --------- Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-authored-by: YuBaoku <49938469+EmmonsCurse@users.noreply.github.com>	2025-11-20 20:00:29 +08:00
周周周	385fe6dade	[Others] clean code (#5133 )	2025-11-20 18:44:08 +08:00
周周周	6fa34102e8	[Others]get_block_shape_and_split_kv_block clean code (#5123 )	2025-11-20 16:40:04 +08:00
freeliuzc	f1e36ff2f7	[Speculative Decoding][MTP]Support stop_seqs and pd-split mode (#5029 ) * support multi_stop_seqs in speculative decoding * support mtp tp with ep split * fix custom op register * fix spec stop_seqs params	2025-11-20 15:26:01 +08:00
chen	9ff418db73	check METAX_GPU (#5114 )	2025-11-19 16:02:21 +08:00
lizhenyun01	d11235333e	format flash_mask_attn	2025-11-18 17:18:12 +08:00
lizhenyun01	cd2c4df64a	format flash_mask_attn	2025-11-18 17:18:12 +08:00
yzwu	d5d0602859	[Iluvatar][CI] disable compiling cudaLaunch API (#5100 )	2025-11-18 14:15:31 +08:00
chen	d58c1db8a0	[Feature][OP] Append Attn Support CUDA-PDL (#5072 )	2025-11-17 20:47:33 +08:00
周周周	b23e684b67	revert group size 3 (#5079 )	2025-11-17 18:54:13 +08:00
Sunny-bot1	8a4ddb29df	Revert "[BugFix] Revert skip capture (#5023 )" (#5080 )	2025-11-17 16:14:55 +08:00
yangjianfengo1	3afb717995	【Fix】fix deepep dispatch (#5036 ) * fix dispatch * fix dispatch --------- Co-authored-by: yuanxiaolan <yuanxiaolan01@baidu.com>	2025-11-17 10:34:01 +08:00
Sunny-bot1	249feca65a	[BugFix] Revert skip capture (#5023 ) Some checks failed CE Compile Job / ce_job_pre_check (push) Has been cancelled Details CE Compile Job / print_ce_job_pre_check_outputs (push) Has been cancelled Details CE Compile Job / FD-Clone-Linux (push) Has been cancelled Details CE Compile Job / Show Code Archive Output (push) Has been cancelled Details CE Compile Job / BUILD_SM8090 (push) Has been cancelled Details CE Compile Job / BUILD_SM8689 (push) Has been cancelled Details CE Compile Job / CE_UPLOAD (push) Has been cancelled Details Deploy GitHub Pages / deploy (push) Has been cancelled Details * Revert "[BugFix][Metax] Fix metax compile issue in get_block_shape_and_split_kv_block (#5000)" This reverts commit `05da8e34c0`. * Revert "skip DtoH capture (#4988)" This reverts commit `5b24013d46`.	2025-11-13 23:52:51 -08:00
周周周	c0a4393d72	[ATTENTION] unitest (#4962 )	2025-11-14 13:45:53 +08:00
carryyu	6c3d1da62f	fix conflicts	2025-11-13 20:30:29 +08:00

1 2 3 4 5

218 Commits