[Cherry-Pick] [BugFix] fix scheduler hang when input length is very close to max_model_len (#5394)
Some checks failed
CE Compile Job / ce_job_pre_check (push) Has been cancelled
CE Compile Job / print_ce_job_pre_check_outputs (push) Has been cancelled
CE Compile Job / FD-Clone-Linux (push) Has been cancelled
CE Compile Job / Show Code Archive Output (push) Has been cancelled
CE Compile Job / BUILD_SM8090 (push) Has been cancelled
CE Compile Job / BUILD_SM8689 (push) Has been cancelled
CE Compile Job / CE_UPLOAD (push) Has been cancelled

* [fix] fix scheduler hang when input length is very close to max_model_len

* [fix] update local_scheduler for v1 scheduler

* [fix] code style
This commit is contained in:
Yonghua Li
2025-12-05 21:51:59 +08:00
committed by GitHub
parent bce3739a57
commit 6961130e04
2 changed files with 4 additions and 3 deletions

View File

@@ -574,7 +574,7 @@ class EngineSevice:
tasks = self.scheduler.get_requests(
available_blocks=available_blocks,
block_size=self.cfg.cache_config.block_size,
reserved_output_blocks=self.cfg.cache_config.enc_dec_block_num,
reserved_output_blocks=0, # self.cfg.cache_config.enc_dec_block_num,
max_num_batched_tokens=self.cfg.max_model_len,
batch=num_prefill_batch,
)