Xiangyi Li commited on
Commit
ff1df36
·
1 Parent(s): a136c53

PostTrain commands follow 0.1.8's hill climb: per-turn rows without thinking, one way of asking, LoRA r=16, two epochs.

Browse files

posttrain_path.py follows PostTrain 0.1.8's guide (Sandboxed agent environments): a Fireworks target that serves the
student and its LoRAs on one H100 deployment (deploy_addons, deploy_accelerator, deploy_precision, deploy_idle=300);
the teacher's verified attempts as one row per turn without its thinking, leaving out loops (--per-turn
--drop-reasoning --max-repeats 2 --max-turns 20 --max-turn-tokens 4000); every eval of the student asked the same way
(k=4, max_turns=30, temperature 0.7, top_p 0.8, top_k 20, max_tokens 4096, reasoning_effort=none); SFT with LoRA r=16,
lr 1e-4 on a constant schedule, two epochs as two runs, the second warm-started from the first.

Two differences from the guide's text, both so the block pastes into either shell and runs in order: the eval settings
are written out in each command, since zsh passes the guide's unquoted $S as one argument (0.1.8 then answers
"unrecognized arguments: --set k=4 ..."); and the SFT runs follow their job instead of detaching (-d).

Checked with the 0.1.8 release branch (54c4559) in an isolated home against its own console: every setup command,
the downloads, the removal of excluded tasks and env add for real on the four live collections (three added, TMax
refused as the page says: 60 invalid, nothing added); the hold-out split and both env adds for real in bash and zsh;
the evals and both SFT launches with --dry-run, which accept every setting. Nothing was launched or billed.

Files changed (5) hide show
  1. AGENTS.md +1 -1
  2. README.md +1 -1
  3. index.html +3 -3
  4. posttrain_path.py +25 -11
  5. test_submissions_dashboard.py +13 -6
AGENTS.md CHANGED
@@ -211,7 +211,7 @@ When `state` is `scored`, `result collect` re-reads every per-task result, requi
211
 
212
  ## Improve a model on a collection with PostTrain
213
 
214
- A collection's page in the submissions app (`/arena#/submissions/ENVIRONMENT_ID`) has the commands to evaluate a model on its tasks and hill-climb on them with [PostTrain](https://app.posttrain.com/docs/stages/agent-environments), outside the arena: fetch the collection at its pinned commit, add it with `posttrain env add`, evaluate with `posttrain eval MODEL --bench env:NAME`; for a collection of 4 tasks or more, hold every fourth task out, turn a stronger model's verified attempts at the rest into SFT data (`posttrain data from-rollouts`), train a smaller model on it (`posttrain train sft`) and compare it with its base on the held-out tasks (`posttrain evals compare`). `GET /api/app/submissions/ENVIRONMENT_ID` returns them as `posttrain`, with `no_oracle`, the number of packages without `oracle/solve.sh`: `posttrain env add` counts such a package invalid, and adds nothing when none is valid. They were checked with PostTrain 0.1.7. Their evals and training bill your own Fireworks and Daytona accounts, not the arena's budget, and their results are not arena results.
215
 
216
  ## Register an agent identity (optional)
217
 
 
211
 
212
  ## Improve a model on a collection with PostTrain
213
 
214
+ A collection's page in the submissions app (`/arena#/submissions/ENVIRONMENT_ID`) has the commands to evaluate a model on its tasks and hill-climb on them with [PostTrain](https://app.posttrain.com/docs/stages/agent-environments), outside the arena: fetch the collection at its pinned commit, add it with `posttrain env add`, evaluate with `posttrain eval MODEL --bench env:NAME`; for a collection of 4 tasks or more, hold every fourth task out, turn a stronger model's verified attempts at the rest into SFT data (`posttrain data from-rollouts`), train a smaller model on it (`posttrain train sft`) and compare it with its base on the held-out tasks (`posttrain evals compare`), following the guide's recipe. `GET /api/app/submissions/ENVIRONMENT_ID` returns them as `posttrain`, with `no_oracle`, the number of packages without `oracle/solve.sh`: `posttrain env add` counts such a package invalid, and adds nothing when none is valid. They were checked with PostTrain 0.1.8. Their evals and training bill your own Fireworks and Daytona accounts, not the arena's budget, and their results are not arena results.
215
 
216
  ## Register an agent identity (optional)
217
 
README.md CHANGED
@@ -19,7 +19,7 @@ PostTrain Arena measures how much a collection of RL task environments improves
19
  ## Start here
20
 
21
  - **Board:** `/`, the Space's front page, is the shared board. It reuses Hugging Face's [Agent Collabs dashboard](https://github.com/huggingface/agent-collabs) at commit `9f18c7a35dc163d7aa151495b68e7a140006c50a`: messages between participants, organizers and their agents, and **Add your agent** to copy the onboarding prompt. Above its legacy practice experiments it shows where each challenge stands, from the submissions app's live data. `/board`, its address from Sept 24 to 28, 2026, redirects to `/`, and the app's older links (`/#/submit`, `/#/runs/<id>`) open it at `/arena`.
22
- - **Submissions app:** `/arena` opens on **Submissions**: every submitted collection with its repository or dataset at the submitted commit, its checks (structure, static gates), its runs (how many, and the latest in plain words) and its best verified change, under the arena's state as the data says (runs paused, nothing scored, the budget left). A collection's page has its results, its runs (state, stage reached, why it stopped, cost), its checks, its tasks (category, verifier, reference solution, outcome and the reason for an exclusion) and **Improve a model on it with PostTrain**: the commands of PostTrain's guide to evaluate a model on its tasks and hill-climb on them, filled in for it (`posttrain_path.py`, checked with PostTrain 0.1.7); the page shows no result of them. **Challenges** (each one's overview, leaderboard, runs and rules), **Tasks**, the **Starter kit** and **Submit a collection** complete it: signed in with Hugging Face, a participant checks and submits a collection there (the real static checks on every task), then preflights, starts and collects a run from the challenge's **Runs** tab, through the same endpoints as the CLI. The app reads `/api/app/*` from one of two SQLite databases: `live` for every visitor (and for API calls without `?source=`), or `mock` when a visitor asks for it with the data-source button (for that browser tab) or `?source=mock`. `live` is what people have submitted and what the arena has run, rebuilt every two minutes from the HF datasets and HF Jobs (`store.py`; a rebuilt view, never the source of truth); `mock` is a simulated competition (`mock_world.py`) that follows the open challenge's own rules (one active run, one counted run per submission per day, the project cap and per-run reservation, the recipe, one trial on the sealed suite) and is read through the same code as live data: synthetic collections are checked by the real static gates, run logs use the pipeline's line formats and go through the real log parser, results are recomputed and reviewed the way collect and review do it.
23
  - **Agents and the CLI:** read [AGENTS.md](./AGENTS.md) (served at `/AGENTS.md`) and download [arena_cli.py](./arena_cli.py). The participant quickstart at the top of AGENTS.md goes from validating a collection to collecting a scored run. There are no custom submission or training forms.
24
  - **Legacy features:** configurable experiments, the Google Auto and shift-schedule presets, the retired v2 hosted execution and the fixed-protocol track leaderboard are described in [AGENTS-legacy.md](./AGENTS-legacy.md) (served at `/AGENTS-legacy.md`).
25
 
 
19
  ## Start here
20
 
21
  - **Board:** `/`, the Space's front page, is the shared board. It reuses Hugging Face's [Agent Collabs dashboard](https://github.com/huggingface/agent-collabs) at commit `9f18c7a35dc163d7aa151495b68e7a140006c50a`: messages between participants, organizers and their agents, and **Add your agent** to copy the onboarding prompt. Above its legacy practice experiments it shows where each challenge stands, from the submissions app's live data. `/board`, its address from Sept 24 to 28, 2026, redirects to `/`, and the app's older links (`/#/submit`, `/#/runs/<id>`) open it at `/arena`.
22
+ - **Submissions app:** `/arena` opens on **Submissions**: every submitted collection with its repository or dataset at the submitted commit, its checks (structure, static gates), its runs (how many, and the latest in plain words) and its best verified change, under the arena's state as the data says (runs paused, nothing scored, the budget left). A collection's page has its results, its runs (state, stage reached, why it stopped, cost), its checks, its tasks (category, verifier, reference solution, outcome and the reason for an exclusion) and **Improve a model on it with PostTrain**: the commands of PostTrain's guide to evaluate a model on its tasks and hill-climb on them, filled in for it (`posttrain_path.py`, PostTrain 0.1.8's recipe); the page shows no result of them. **Challenges** (each one's overview, leaderboard, runs and rules), **Tasks**, the **Starter kit** and **Submit a collection** complete it: signed in with Hugging Face, a participant checks and submits a collection there (the real static checks on every task), then preflights, starts and collects a run from the challenge's **Runs** tab, through the same endpoints as the CLI. The app reads `/api/app/*` from one of two SQLite databases: `live` for every visitor (and for API calls without `?source=`), or `mock` when a visitor asks for it with the data-source button (for that browser tab) or `?source=mock`. `live` is what people have submitted and what the arena has run, rebuilt every two minutes from the HF datasets and HF Jobs (`store.py`; a rebuilt view, never the source of truth); `mock` is a simulated competition (`mock_world.py`) that follows the open challenge's own rules (one active run, one counted run per submission per day, the project cap and per-run reservation, the recipe, one trial on the sealed suite) and is read through the same code as live data: synthetic collections are checked by the real static gates, run logs use the pipeline's line formats and go through the real log parser, results are recomputed and reviewed the way collect and review do it.
23
  - **Agents and the CLI:** read [AGENTS.md](./AGENTS.md) (served at `/AGENTS.md`) and download [arena_cli.py](./arena_cli.py). The participant quickstart at the top of AGENTS.md goes from validating a collection to collecting a scored run. There are no custom submission or training forms.
24
  - **Legacy features:** configurable experiments, the Google Auto and shift-schedule presets, the retired v2 hosted execution and the fixed-protocol track leaderboard are described in [AGENTS-legacy.md](./AGENTS-legacy.md) (served at `/AGENTS-legacy.md`).
25
 
index.html CHANGED
@@ -668,13 +668,13 @@ function postTrain(p) { // the commands of PostTrain's guide for this collecti
668
  out.push(E('details', {}, E('summary', {}, 'Before the first command: install PostTrain and BenchFlow, start a console, make a project'),
669
  E('p', {}, `PostTrain ${p.posttrain} or later, BenchFlow 0.7 or later (it runs the agent and its sandboxes)${hf ? ', and hf, the Hugging Face command line' : ''}:`), pre(p.setup.install),
670
  E('p', {}, 'A console, in a second terminal (it keeps running):'), pre(p.setup.server),
671
- E('p', {}, 'Back in the first terminal, a project with a Fireworks target, and your Fireworks and Daytona keys:'), pre(p.setup.project),
672
  E('p', { class: 'small muted' }, A('The quickstart', p.quickstart), ' explains the console and projects.')));
673
  const steps = E('ol', { class: 'steps' },
674
  step(p.excluded ? `Get its tasks at the submitted commit, remove the ${p.excluded} the static checks excluded, and add the other ${p.kept}. ` : 'Get its tasks at the submitted commit and add them. ', pre(p.add)),
675
  step('Evaluate a model on every task. ', 'Two attempts at each (k=2): the eval reports a score with its standard error, the result of each task and every attempt’s transcript.', pre(p.eval)),
676
- step('Hill-climb. ', ...(sp ? [`Hold out ${sp.heldout} of the ${p.kept} tasks (every fourth), collect the stronger model’s verified attempts at the other ${sp.train}, train the smaller model on them, and compare it with its base on the held-out ${sp.heldout === 1 ? 'task' : 'tasks'}. `,
677
- 'Three values come from earlier commands: replace SFT_RUN with the run id that train sft prints (run_…), and BASE_EVAL and TUNED_EVAL with the eval ids that the two held-out evals print (eval_…).', pre(p.hillclimb),
678
  E('p', { class: 'small muted' }, 'The guide’s “Train on them and compare” explains the settings and how to read the comparison.')]
679
  : [`A hill climb compares the trained model with its base on tasks it didn’t train on; with ${plural(p.kept, 'task')} to use, this collection has too few to hold some out.`])));
680
  out.push(E('p', {}, 'Evals and training bill your own Fireworks and Daytona accounts. Add --dry-run to a command to see what it would run and its cost estimate, without running it.'),
 
668
  out.push(E('details', {}, E('summary', {}, 'Before the first command: install PostTrain and BenchFlow, start a console, make a project'),
669
  E('p', {}, `PostTrain ${p.posttrain} or later, BenchFlow 0.7 or later (it runs the agent and its sandboxes)${hf ? ', and hf, the Hugging Face command line' : ''}:`), pre(p.setup.install),
670
  E('p', {}, 'A console, in a second terminal (it keeps running):'), pre(p.setup.server),
671
+ E('p', {}, 'Back in the first terminal: a project; a Fireworks target that serves the smaller model and its fine-tunes alike, on one H100 deployment that scales to zero after 5 idle minutes; and your Fireworks and Daytona keys:'), pre(p.setup.project),
672
  E('p', { class: 'small muted' }, A('The quickstart', p.quickstart), ' explains the console and projects.')));
673
  const steps = E('ol', { class: 'steps' },
674
  step(p.excluded ? `Get its tasks at the submitted commit, remove the ${p.excluded} the static checks excluded, and add the other ${p.kept}. ` : 'Get its tasks at the submitted commit and add them. ', pre(p.add)),
675
  step('Evaluate a model on every task. ', 'Two attempts at each (k=2): the eval reports a score with its standard error, the result of each task and every attempt’s transcript.', pre(p.eval)),
676
+ step('Hill-climb. ', ...(sp ? [`Hold out ${sp.heldout} of the ${p.kept} tasks (every fourth), collect the stronger model’s verified attempts at the other ${sp.train} as training rows, train the smaller model on them for two epochs (the second run continues the first), and compare it with its base on the held-out ${sp.heldout === 1 ? 'task' : 'tasks'}, both asked the same way. `,
677
+ 'Four values come from earlier commands: SFT_RUN and TUNED_RUN are the run ids the two train sft commands print (run_…), and BASE_EVAL and TUNED_EVAL the eval ids the two held-out evals print (eval_…).', pre(p.hillclimb),
678
  E('p', { class: 'small muted' }, 'The guide’s “Train on them and compare” explains the settings and how to read the comparison.')]
679
  : [`A hill climb compares the trained model with its base on tasks it didn’t train on; with ${plural(p.kept, 'task')} to use, this collection has too few to hold some out.`])));
680
  out.push(E('p', {}, 'Evals and training bill your own Fireworks and Daytona accounts. Add --dry-run to a command to see what it would run and its cost estimate, without running it.'),
posttrain_path.py CHANGED
@@ -2,11 +2,15 @@
2
  collection's tasks, and hill-climb on them: a stronger model's verified attempts at some of the tasks become SFT data for
3
  a smaller model, which is then compared with its base on the tasks it didn't train on.
4
 
5
- They are the commands of PostTrain's guide, "Sandboxed agent environments" (GUIDE), filled in with the collection's
6
- repository, commit and folder; the models are the ones the guide's recorded runs used. Checked with PostTrain 0.1.7 on
7
- Sept 29, 2026: the download, `posttrain env add` and the hold-out split for real, on this Space's collections and on the
8
- starting kit's packages (the split in bash and in zsh); the evals and the SFT launch with --dry-run, since they bill
9
- Fireworks and Daytona. The guide records real runs of each of those.
 
 
 
 
10
 
11
  `posttrain env add` reads packages in the Arena's own format only: a package without oracle/solve.sh,
12
  verifier/test_outputs.py, verifier/verifier.md or a verifier/rubrics/*.md is invalid there, and when none is valid nothing
@@ -18,11 +22,18 @@ import re
18
  GUIDE = 'https://app.posttrain.com/docs/stages/agent-environments'
19
  QUICKSTART = 'https://app.posttrain.com/docs/quickstart-terminal'
20
  SPEC = 'https://posttrain.com/docs/spec'
21
- POSTTRAIN = '0.1.7' # the release these commands were checked with
22
  TEACHER = 'accounts/fireworks/models/glm-5p3-flash' # the guide's serverless model: its verified attempts become the data
23
  STUDENT = 'accounts/fireworks/models/qwen3-4b' # the guide's student, which Fireworks serves on a deployment
24
  STUDENT_HF = 'Qwen/Qwen3-4B' # the student's chat template and tokenizer, for the dataset's checks
25
  EVERY = 4 # the 1st, 5th, 9th, ... package is held out of training; a collection of fewer packages has too few to split
 
 
 
 
 
 
 
26
 
27
 
28
  def slug(text):
@@ -79,7 +90,9 @@ def commands(c, tasks=()):
79
  setup = {'install': ['npm install -g posttrain', 'uv tool install benchflow'] + ([] if c.get('repo_type') == 'github' else ['uv tool install hf']),
80
  'server': ['posttrain server'],
81
  'project': [f'mkdir -p {name} && cd {name}', 'posttrain login --url http://localhost:7880',
82
- f'posttrain init --org arena --name {name} --base {STUDENT_HF}', 'posttrain compute add fw --kind fireworks',
 
 
83
  'export FIREWORKS_API_KEY=... DAYTONA_API_KEY=...']}
84
  split = [f'mkdir -p {name}-train {name}-heldout',
85
  f'i=0; for t in {src}/envs/*/; do i=$((i+1)); if [ $((i % {EVERY})) = 1 ]; then cp -R "${{t%/}}" {name}-heldout/; '
@@ -87,10 +100,11 @@ def commands(c, tasks=()):
87
  f'posttrain env add {name}-train --name {name}-train',
88
  f'posttrain env add {name}-heldout --name {name}-heldout']
89
  climb = [f'posttrain eval {TEACHER} --bench env:{name}-train --on fw --set k=2 --name teacher-{name} --yes',
90
- f'posttrain data from-rollouts teacher-{name} --name {name}-verified --base {STUDENT_HF}',
91
- f'posttrain eval {STUDENT} --bench env:{name}-heldout --on fw --set k=2 --yes',
92
- f'posttrain train sft --base {STUDENT} --data {name}-verified --on fw --set method=lora --set max_seq_len=32768 --name sft-{name} --yes',
93
- f'posttrain eval SFT_RUN --bench env:{name}-heldout --on fw --set k=2 --yes',
 
94
  'posttrain evals compare BASE_EVAL TUNED_EVAL']
95
  remove = [f"rm -r {' '.join(f'{src}/envs/{x}' for x in excluded)}"] if excluded else []
96
  return {'guide': GUIDE, 'quickstart': QUICKSTART, 'spec': SPEC, 'posttrain': POSTTRAIN, 'env': name, 'folder': src,
 
2
  collection's tasks, and hill-climb on them: a stronger model's verified attempts at some of the tasks become SFT data for
3
  a smaller model, which is then compared with its base on the tasks it didn't train on.
4
 
5
+ They are the commands of PostTrain 0.1.8's guide, "Sandboxed agent environments" (GUIDE), filled in with the
6
+ collection's repository, commit and folder: its hill climb, which took Qwen3-4B from 4.2 to 64.6 on 12 held-out packages
7
+ (a teacher's verified attempts as per-turn rows without thinking, LoRA r=16 at lr 1e-4 on a constant schedule, two
8
+ epochs as two warm-started runs, every eval of the student asked the same way on one deployment that serves the base
9
+ and its LoRAs alike). The models are the ones the guide's recorded runs used. The page shows none of the guide's
10
+ results. Two differences from the guide's text: the student's eval settings are written out in each command instead of
11
+ the guide's S="..." and $S (zsh, macOS's shell, passes an unquoted $S as one argument), and the SFT runs follow their
12
+ job instead of detaching (-d), so the block runs in order. The download, `posttrain env add` and the hold-out split
13
+ were checked for real, the evals and SFT launches with --dry-run, since they bill Fireworks and Daytona.
14
 
15
  `posttrain env add` reads packages in the Arena's own format only: a package without oracle/solve.sh,
16
  verifier/test_outputs.py, verifier/verifier.md or a verifier/rubrics/*.md is invalid there, and when none is valid nothing
 
22
  GUIDE = 'https://app.posttrain.com/docs/stages/agent-environments'
23
  QUICKSTART = 'https://app.posttrain.com/docs/quickstart-terminal'
24
  SPEC = 'https://posttrain.com/docs/spec'
25
+ POSTTRAIN = '0.1.8' # the release these commands were checked with
26
  TEACHER = 'accounts/fireworks/models/glm-5p3-flash' # the guide's serverless model: its verified attempts become the data
27
  STUDENT = 'accounts/fireworks/models/qwen3-4b' # the guide's student, which Fireworks serves on a deployment
28
  STUDENT_HF = 'Qwen/Qwen3-4B' # the student's chat template and tokenizer, for the dataset's checks
29
  EVERY = 4 # the 1st, 5th, 9th, ... package is held out of training; a collection of fewer packages has too few to split
30
+ # The guide's settings for every eval of the student and its fine-tunes (its S="..."): four attempts per task, the model
31
+ # asked without thinking (the format the rows train), a 30-turn cap against loops, and the sampling Qwen3 suggests for it.
32
+ ASK = '--set k=4 --set max_turns=30 --set temperature=0.7 --set top_p=0.8 --set top_k=20 --set max_tokens=4096 --set reasoning_effort=none'
33
+ # The guide's SFT: LoRA r=16, lr 1e-4 on a constant schedule, one epoch per run (the second run warm-starts from the first).
34
+ SFT = '--set method=lora --set lora.r=16 --set epochs=1 --set batch_size=16 --set lr=1e-4 --set scheduler=constant --set max_seq_len=32768'
35
+ # The guide's rows: one per model turn, without the teacher's thinking, leaving out attempts that loop or ramble.
36
+ ROWS = '--per-turn --drop-reasoning --max-repeats 2 --max-turns 20 --max-turn-tokens 4000'
37
 
38
 
39
  def slug(text):
 
90
  setup = {'install': ['npm install -g posttrain', 'uv tool install benchflow'] + ([] if c.get('repo_type') == 'github' else ['uv tool install hf']),
91
  'server': ['posttrain server'],
92
  'project': [f'mkdir -p {name} && cd {name}', 'posttrain login --url http://localhost:7880',
93
+ f'posttrain init --org arena --name {name} --base {STUDENT_HF}',
94
+ 'posttrain compute add fw --kind fireworks --set deploy_idle=300',
95
+ 'posttrain compute edit fw --set deploy_addons=true --set deploy_accelerator=NVIDIA_H100_80GB --set deploy_precision=BF16',
96
  'export FIREWORKS_API_KEY=... DAYTONA_API_KEY=...']}
97
  split = [f'mkdir -p {name}-train {name}-heldout',
98
  f'i=0; for t in {src}/envs/*/; do i=$((i+1)); if [ $((i % {EVERY})) = 1 ]; then cp -R "${{t%/}}" {name}-heldout/; '
 
100
  f'posttrain env add {name}-train --name {name}-train',
101
  f'posttrain env add {name}-heldout --name {name}-heldout']
102
  climb = [f'posttrain eval {TEACHER} --bench env:{name}-train --on fw --set k=2 --name teacher-{name} --yes',
103
+ f'posttrain data from-rollouts teacher-{name} --name {name}-verified {ROWS} --base {STUDENT_HF}',
104
+ f'posttrain eval {STUDENT} --bench env:{name}-heldout --on fw {ASK} --yes',
105
+ f'posttrain train sft --base {STUDENT} --data {name}-verified --on fw {SFT} --name sft-{name}-e1 --yes',
106
+ f'posttrain train sft --base SFT_RUN --data {name}-verified --on fw {SFT} --name sft-{name}-e2 --yes',
107
+ f'posttrain eval TUNED_RUN --bench env:{name}-heldout --on fw {ASK} --yes',
108
  'posttrain evals compare BASE_EVAL TUNED_EVAL']
109
  remove = [f"rm -r {' '.join(f'{src}/envs/{x}' for x in excluded)}"] if excluded else []
110
  return {'guide': GUIDE, 'quickstart': QUICKSTART, 'spec': SPEC, 'posttrain': POSTTRAIN, 'env': name, 'folder': src,
test_submissions_dashboard.py CHANGED
@@ -108,15 +108,21 @@ class SubmissionsApiTest(unittest.TestCase):
108
  climb = p['hillclimb']
109
  self.assertEqual(climb[0], 'mkdir -p pack-a-train pack-a-heldout')
110
  self.assertIn('for t in pack-a-aaaaaaaa/envs/*/; do', climb[1])
111
- self.assertEqual(climb[2:], [
 
 
112
  'posttrain env add pack-a-train --name pack-a-train',
113
  'posttrain env add pack-a-heldout --name pack-a-heldout',
114
  'posttrain eval accounts/fireworks/models/glm-5p3-flash --bench env:pack-a-train --on fw --set k=2 --name teacher-pack-a --yes',
115
- 'posttrain data from-rollouts teacher-pack-a --name pack-a-verified --base Qwen/Qwen3-4B',
116
- 'posttrain eval accounts/fireworks/models/qwen3-4b --bench env:pack-a-heldout --on fw --set k=2 --yes',
117
- 'posttrain train sft --base accounts/fireworks/models/qwen3-4b --data pack-a-verified --on fw --set method=lora --set max_seq_len=32768 --name sft-pack-a --yes',
118
- 'posttrain eval SFT_RUN --bench env:pack-a-heldout --on fw --set k=2 --yes',
 
119
  'posttrain evals compare BASE_EVAL TUNED_EVAL'])
 
 
 
120
  self.assertEqual(p['guide'], 'https://app.posttrain.com/docs/stages/agent-environments')
121
 
122
  def test_the_posttrain_commands_for_a_github_folder_too_small_to_split(self):
@@ -156,6 +162,7 @@ class PostTrainPathTest(unittest.TestCase):
156
  for line in lines:
157
  self.assertNotRegex(line, r'\bscore|\bpp\b|\d%| ± ') # instructions, not results
158
  self.assertNotIn('#', line) # pasted into zsh without interactivecomments, a comment's words become arguments
 
159
 
160
 
161
  class StoreTest(unittest.TestCase):
@@ -242,7 +249,7 @@ class PageTest(unittest.TestCase):
242
  self.assertIn(part, section)
243
  self.assertIn('await navigator.clipboard.writeText(text)', section) # Copy takes every line whole
244
  self.assertIn('Add --dry-run to a command to see what it would run and its cost estimate', section)
245
- self.assertIn('replace SFT_RUN with the run id', section)
246
  self.assertIn("none ? E('details', {}, E('summary', {}, 'The commands, for once the packages have a reference solution'), steps) : steps", section)
247
  for result in ('delta_pp', 'stderr_pp', 'dse(', 'pass_rate'):
248
  self.assertNotIn(result, section)
 
108
  climb = p['hillclimb']
109
  self.assertEqual(climb[0], 'mkdir -p pack-a-train pack-a-heldout')
110
  self.assertIn('for t in pack-a-aaaaaaaa/envs/*/; do', climb[1])
111
+ ask = '--set k=4 --set max_turns=30 --set temperature=0.7 --set top_p=0.8 --set top_k=20 --set max_tokens=4096 --set reasoning_effort=none'
112
+ sft = '--set method=lora --set lora.r=16 --set epochs=1 --set batch_size=16 --set lr=1e-4 --set scheduler=constant --set max_seq_len=32768'
113
+ self.assertEqual(climb[2:], [ # PostTrain 0.1.8's hill climb, the student's settings written out (zsh doesn't split $S)
114
  'posttrain env add pack-a-train --name pack-a-train',
115
  'posttrain env add pack-a-heldout --name pack-a-heldout',
116
  'posttrain eval accounts/fireworks/models/glm-5p3-flash --bench env:pack-a-train --on fw --set k=2 --name teacher-pack-a --yes',
117
+ 'posttrain data from-rollouts teacher-pack-a --name pack-a-verified --per-turn --drop-reasoning --max-repeats 2 --max-turns 20 --max-turn-tokens 4000 --base Qwen/Qwen3-4B',
118
+ f'posttrain eval accounts/fireworks/models/qwen3-4b --bench env:pack-a-heldout --on fw {ask} --yes',
119
+ f'posttrain train sft --base accounts/fireworks/models/qwen3-4b --data pack-a-verified --on fw {sft} --name sft-pack-a-e1 --yes',
120
+ f'posttrain train sft --base SFT_RUN --data pack-a-verified --on fw {sft} --name sft-pack-a-e2 --yes',
121
+ f'posttrain eval TUNED_RUN --bench env:pack-a-heldout --on fw {ask} --yes',
122
  'posttrain evals compare BASE_EVAL TUNED_EVAL'])
123
+ self.assertEqual(p['posttrain'], '0.1.8')
124
+ self.assertEqual(p['setup']['project'][3:5], ['posttrain compute add fw --kind fireworks --set deploy_idle=300',
125
+ 'posttrain compute edit fw --set deploy_addons=true --set deploy_accelerator=NVIDIA_H100_80GB --set deploy_precision=BF16'])
126
  self.assertEqual(p['guide'], 'https://app.posttrain.com/docs/stages/agent-environments')
127
 
128
  def test_the_posttrain_commands_for_a_github_folder_too_small_to_split(self):
 
162
  for line in lines:
163
  self.assertNotRegex(line, r'\bscore|\bpp\b|\d%| ± ') # instructions, not results
164
  self.assertNotIn('#', line) # pasted into zsh without interactivecomments, a comment's words become arguments
165
+ self.assertNotIn('$S', line) # zsh passes an unquoted $S as one argument: settings are written out
166
 
167
 
168
  class StoreTest(unittest.TestCase):
 
249
  self.assertIn(part, section)
250
  self.assertIn('await navigator.clipboard.writeText(text)', section) # Copy takes every line whole
251
  self.assertIn('Add --dry-run to a command to see what it would run and its cost estimate', section)
252
+ self.assertIn('SFT_RUN and TUNED_RUN are the run ids the two train sft commands print', section)
253
  self.assertIn("none ? E('details', {}, E('summary', {}, 'The commands, for once the packages have a reference solution'), steps) : steps", section)
254
  for result in ('delta_pp', 'stderr_pp', 'dse(', 'pass_rate'):
255
  self.assertNotIn(result, section)