|
Download README.md from novastar111/sokoban_adaptive_mixed_forward_rl_step50: direct link, hf CLI and curl.
- Browser
- Download file 934 Bytes
-
https://huggingface.co/novastar111/sokoban_adaptive_mixed_forward_rl_step50/resolve/main/README.md
- Command line
-
hf download hf://novastar111/sokoban_adaptive_mixed_forward_rl_step50/README.md
-
curl -L -o README.md https://huggingface.co/novastar111/sokoban_adaptive_mixed_forward_rl_step50/resolve/main/README.md
934 Bytes
metadata
license: cc-by-nc-4.0
tags:
- bagel
- vlm-gym
- sokoban
- reinforcement-learning
- adaptive-thinking
sokoban_adaptive_mixed_forward_rl_step50
BAGEL-7B-MoT Sokoban checkpoint after 50 IMP-agent RL steps.
- arm: mixed-forward
- initialization:
novastar111/sokoban_adaptive_mixed_forward_sft3k - RL data: 2,000 mixed 3-box boards (certified deadlock + trivial)
- objective: environment success with a
-0.1malformed-output penalty - weights: converted BF16 EMA safetensors; optimizer/training state is not included
The mixed-forward and mixed-branch labels refer to the step-3000 SFT initializations used by the
RL experiment. Evaluation uses Sokoban-v8 q95/perseg, full self-rollout, move-only actions, and
stop-required success. Load with the public BAGEL-7B-MoT base/config and point the eval runner's
--bagel-path at this repository.