FastDeploy

mirror of https://github.com/PaddlePaddle/FastDeploy.git synced 2025-12-24 13:28:13 +08:00

Author	SHA1	Message	Date
GoldPancake	e56c4dd0a8	[Cherry-Pick] Support for request-level speculative decoding metrics monitoring.(#5518 ) (#5614 ) * support spec metrics monitor per request	2025-12-17 20:53:04 +08:00
kxz2002	97189079b9	[BugFix] unify max_tokens (#4968 ) * unify max tokens * modify and add unit test * modify and add unit test * modify and add unit tests --------- Co-authored-by: YuBaoku <49938469+EmmonsCurse@users.noreply.github.com>	2025-11-18 20:01:33 +08:00
LiqinruiG	4251ac5e95	【Fix】 remove text_after_process & raw_prediction (#4421 ) * remove text_after_process & raw_prediction * remove text_after_process & raw_prediction	2025-10-16 19:00:18 +08:00
zhuzixuan	a47976e82d	[Echo] Support more types of prompt echo (#4022 ) * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. * wenxin-tools-700 When the prompt type is list[int] or list[list[int]], it needs to support echoing after decoding. --------- Co-authored-by: luukunn <83932082+luukunn@users.noreply.github.com>	2025-09-11 19:34:44 +08:00
SunLei	b9af95cf1c	[Feature] Add AsyncTokenizerClient&ChatResponseProcessor with remote encode&decode support. (#3674 ) * [Feature] add AsyncTokenizerClient * add decode_image * Add response_processors with remote decode support. * [Feature] add tokenizer_base_url startup argument * Revert comment removal and restore original content. * [Feature] Non-streaming requests now support remote image decoding. * Fix parameter type issue in decode_image call. * Keep completion_token_ids when return_token_ids = False. * add copyright	2025-08-30 17:06:26 +08:00
Yzc216	466cbb5a99	[Feature] Models api (#3073 ) * add v1/models interface related * add model parameters * default model verification * unit test * check model err_msg * unit test * type annotation * model parameter in response * modify document description * modify document description * unit test * verification * verification update * model_name * pre-commit * update test case * update test case * Update tests/entrypoints/openai/test_serving_models.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update tests/entrypoints/openai/test_serving_models.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update tests/entrypoints/openai/test_serving_models.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update tests/entrypoints/openai/test_serving_models.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update fastdeploy/entrypoints/openai/serving_models.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: LiqinruiG <37392159+LiqinruiG@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>	2025-08-21 17:02:56 +08:00
YUNSHEN XIE	3a6058e445	Add stable ci (#3460 ) * add stable ci * fix * update * fix * rename tests dir;fix stable ci bug * add timeout limit * update	2025-08-20 08:57:17 +08:00

7 Commits