gaoziyuan
|
a592d17615
|
support qwen3 name_mapping (#3180)
|
2025-08-04 16:37:34 +08:00 |
|
Zero Rains
|
0fb37ab7e4
|
update flake8 version to support pre-commit in python3.12 (#3000)
* update flake8 version to support pre-commit in python3.12
* polish code
|
2025-07-24 01:43:31 -07:00 |
|
gaoziyuan
|
95a214ae43
|
support trainer_degree in name_mapping (#2935)
|
2025-07-20 23:12:55 -07:00 |
|
Zero Rains
|
25698d56d1
|
polish code with new pre-commit rule (#2923)
|
2025-07-19 23:19:27 +08:00 |
|
gaoziyuan
|
6efad14b95
|
support vl ori_vacab_size (#2900)
|
2025-07-18 16:26:14 +08:00 |
|
Yuanle Liu
|
dbb9e2506b
|
Fix rollout_model init (#2881)
|
2025-07-16 22:36:21 -07:00 |
|
Yuanle Liu
|
61b3997b85
|
refactor rl get_name_mappings_to_training (#2847)
Deploy GitHub Pages / deploy (push) Has been cancelled
* refactor rl get_name_mappings_to_training
* fix tp>1
* change variable name(ffn1->up_gate_proj/ffn2->down_proj)
* change variable name(linear_weight->weight/linear_bias->bias)
* add rl names mapping for vl
* fix ernie 0.3B error
* fix develop code
* fix
|
2025-07-15 07:31:42 -07:00 |
|
YuanRisheng
|
4c7b8bc458
|
Simplify the Config code (#2770)
* simplify the code
* fix vl
* delete config
* fix
* perfect code
* fix ci
* fix xpu
* fix xpu
* fix server
* resolve conflict
* fix mtp
* resolve conflict
* fix xpu
* fix xpu
* fix vl
* fix log
* fix qwen moe
* fix qwen moe
* fix qwen moe
|
2025-07-14 19:50:05 +08:00 |
|
gaoziyuan
|
749b2e9c89
|
support qwen3moe name_mapping (#2820)
|
2025-07-12 12:05:54 +08:00 |
|
gaoziyuan
|
26d5d737dd
|
【Fearture】support qwen2 some func (#2740)
* add rl qwen model support
* fix
* fix
|
2025-07-08 12:03:04 +08:00 |
|
Jiang-Jia-Jun
|
05c670e593
|
[Sync] Update to latest code (#2679)
* [Sync] Update to latest code
* Add new code files
* Add new code files
* update code
* Try to fix build.sh
* Try to fix build.sh
* Update code
* Update requirements.txt
* Update code
---------
Co-authored-by: Jiang-Jia-Jun <jiangjiajun@baidu.com>
|
2025-07-03 15:43:53 +08:00 |
|