[Feature][MTP]Support MTP for rl-model (#4009)

* qk norm for speculate decode C16 * support mtp in v1_scheduler mode * support mtp rope_3d * support mtp features * add unit test && del some log --------- Co-authored-by: yuanxiaolan <yuanxiaolan01@baidu.com> Co-authored-by: xiaoxiaohehe001 <hiteezsf@163.com>
2025-10-05 08:37:06 +08:00 · 2025-09-10 13:34:37 +08:00
parent cce2410fad
commit 2f473ba966
21 changed files with 1465 additions and 531 deletions
--- a/fastdeploy/model_executor/pre_and_post_process.py
+++ b/fastdeploy/model_executor/pre_and_post_process.py
@@ -306,7 +306,9 @@ def post_process_normal(
            )


-def post_process_specualate(model_output, save_each_rank: bool = False, skip_save_output: bool = False):
+def post_process_specualate(
+    model_output: ModelOutputData, save_each_rank: bool = False, skip_save_output: bool = False
+):
    """"""
    speculate_update(
        model_output.seq_lens_encoder,