1202kbs commited on
Commit
c7ade81
·
verified ·
1 Parent(s): c790363

FAR release bundles

Browse files
README.md CHANGED
@@ -54,19 +54,24 @@ or `huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models`.
54
  | AI2-THOR-dyn | `ai2thor_dyn/far_multicue` | FAR -- Multi-Cue | 500k | 0.56 GB |
55
  | AI2-THOR-dyn | `ai2thor_dyn/temporal` | Temporal | 500k | 0.49 GB |
56
  | AI2-THOR-dyn | `ai2thor_dyn/worldmem` | WorldMem | 500k | 0.49 GB |
 
57
  | AI2-THOR v3 | `ai2thor_v3/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
58
  | AI2-THOR v3 | `ai2thor_v3/temporal` | Temporal | 700k | 0.49 GB |
59
  | AI2-THOR v3 | `ai2thor_v3/worldmem` | WorldMem | 700k | 0.49 GB |
 
60
  | LoopNav | `loopnav/far_meta` | FAR -- Meta | 700k | 0.55 GB |
61
  | LoopNav | `loopnav/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
62
  | LoopNav | `loopnav/far_visual` | FAR -- Visual | 700k | 0.55 GB |
63
  | LoopNav | `loopnav/longlive_rag` | LongLive-RAG | 700k | 0.55 GB |
64
  | LoopNav | `loopnav/temporal` | Temporal | 700k | 0.49 GB |
65
  | LoopNav | `loopnav/worldmem` | WorldMem | 700k | 0.49 GB |
 
 
66
  | SoundSpaces v1 | `soundspaces_v1/far_meta` | FAR -- Meta | 700k | 0.55 GB |
67
  | SoundSpaces v1 | `soundspaces_v1/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
68
  | SoundSpaces v1 | `soundspaces_v1/temporal` | Temporal | 700k | 0.49 GB |
69
  | SoundSpaces v1 | `soundspaces_v1/worldmem` | WorldMem | 700k | 0.49 GB |
 
70
  | SoundSpaces v2 | `soundspaces_v2/far_meta` | FAR -- Meta | 700k | 0.55 GB |
71
  | SoundSpaces v2 | `soundspaces_v2/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
72
  | SoundSpaces v2 | `soundspaces_v2/temporal` | Temporal | 700k | 0.49 GB |
@@ -74,7 +79,9 @@ or `huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models`.
74
 
75
  "Temporal" and "WorldMem" are the recency and field-of-view retrieval
76
  baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms
77
- differ in the cues the retriever fuses (metadata, visual, multi-cue).
 
 
78
 
79
  ## License
80
 
 
54
  | AI2-THOR-dyn | `ai2thor_dyn/far_multicue` | FAR -- Multi-Cue | 500k | 0.56 GB |
55
  | AI2-THOR-dyn | `ai2thor_dyn/temporal` | Temporal | 500k | 0.49 GB |
56
  | AI2-THOR-dyn | `ai2thor_dyn/worldmem` | WorldMem | 500k | 0.49 GB |
57
+ | AI2-THOR-dyn | `ai2thor_dyn/retriever_object` | retriever (object cue), init of the FAR arms | 12.5k | 0.06 GB |
58
  | AI2-THOR v3 | `ai2thor_v3/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
59
  | AI2-THOR v3 | `ai2thor_v3/temporal` | Temporal | 700k | 0.49 GB |
60
  | AI2-THOR v3 | `ai2thor_v3/worldmem` | WorldMem | 700k | 0.49 GB |
61
+ | AI2-THOR v3 | `ai2thor_v3/retriever_jepa` | retriever (visual), init of the FAR arm | 400k | 0.06 GB |
62
  | LoopNav | `loopnav/far_meta` | FAR -- Meta | 700k | 0.55 GB |
63
  | LoopNav | `loopnav/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
64
  | LoopNav | `loopnav/far_visual` | FAR -- Visual | 700k | 0.55 GB |
65
  | LoopNav | `loopnav/longlive_rag` | LongLive-RAG | 700k | 0.55 GB |
66
  | LoopNav | `loopnav/temporal` | Temporal | 700k | 0.49 GB |
67
  | LoopNav | `loopnav/worldmem` | WorldMem | 700k | 0.49 GB |
68
+ | LoopNav | `loopnav/retriever_jepa` | retriever (visual), init of the FAR arms | 400k | 0.06 GB |
69
+ | LoopNav | `loopnav/retriever_longlive` | retriever (content query), init of LongLive-RAG | 400k | 0.06 GB |
70
  | SoundSpaces v1 | `soundspaces_v1/far_meta` | FAR -- Meta | 700k | 0.55 GB |
71
  | SoundSpaces v1 | `soundspaces_v1/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
72
  | SoundSpaces v1 | `soundspaces_v1/temporal` | Temporal | 700k | 0.49 GB |
73
  | SoundSpaces v1 | `soundspaces_v1/worldmem` | WorldMem | 700k | 0.49 GB |
74
+ | SoundSpaces v1 | `soundspaces_v1/retriever_audio` | retriever (audio), init of the SoundSpaces v1 and v2 FAR arms | 400k | 0.06 GB |
75
  | SoundSpaces v2 | `soundspaces_v2/far_meta` | FAR -- Meta | 700k | 0.55 GB |
76
  | SoundSpaces v2 | `soundspaces_v2/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
77
  | SoundSpaces v2 | `soundspaces_v2/temporal` | Temporal | 700k | 0.49 GB |
 
79
 
80
  "Temporal" and "WorldMem" are the recency and field-of-view retrieval
81
  baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms
82
+ differ in the cues the retriever fuses (metadata, visual, multi-cue). The retriever bundles are
83
+ the contrastively pre-trained encoders the FAR launchers start from; evaluation does not need
84
+ them (each FAR bundle already carries its trained retriever).
85
 
86
  ## License
87
 
ai2thor_dyn/retriever_object/checkpoints/0012500.pth.tar ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6d3d3e13be18d6e5d75bfb6ae59ef5e4965b11f48ae22133db917a99a44222f5
3
+ size 66429235
ai2thor_dyn/retriever_object/config.yaml ADDED
@@ -0,0 +1,948 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ recon:
2
+ spec:
3
+ _target_: far.builders.data.build_dataspec
4
+ local_pos_stats:
5
+ min:
6
+ - -2.5
7
+ - -4
8
+ max:
9
+ - 5
10
+ - 4
11
+ meters_per_waypoint: 0.25
12
+ max_frame_offset: 128
13
+ len_traj_pred: ${len_traj_pred}
14
+ planner_mu_init:
15
+ - -0.1
16
+ - 0
17
+ - 0
18
+ planner_sigma_init:
19
+ - 0.02
20
+ - 0.1
21
+ - 0.3142
22
+ pose_fields:
23
+ - x
24
+ - 'y'
25
+ - z
26
+ - pitch
27
+ - yaw
28
+ pose_dim: 5
29
+ pos_dim: 2
30
+ angle_dim: 1
31
+ pose_mode: 2d
32
+ fps: 4.0
33
+ shared:
34
+ manifest_path: results/manifests/recon/recon_split.json
35
+ backend:
36
+ _target_: far.data.backends.NavigationBackend
37
+ transform:
38
+ _target_: far.data.transforms.make_centercrop_resize
39
+ image_height: ${image_height}
40
+ image_width: ${image_width}
41
+ sampler:
42
+ context_goal:
43
+ _target_: far.data.samplers.ContextGoalSampler
44
+ context_pool_size: 12
45
+ len_traj_pred: ${len_traj_pred}
46
+ goals_per_obs: ${goals_per_obs}
47
+ context_traj:
48
+ _target_: far.data.samplers.ContextTrajectorySampler
49
+ context_pool_size: 12
50
+ len_traj_pred: ${len_traj_pred}
51
+ formatter:
52
+ context_pose:
53
+ _target_: far.data.formatters.ContextPoseFormatter2D
54
+ spec: ${recon.spec}
55
+ train:
56
+ _target_: far.data.datasets.FormattedVideoDataset
57
+ formatter: ${recon.formatter.context_pose}
58
+ base_dataset:
59
+ _target_: far.data.datasets.VideoDataset
60
+ manifest_path: ${recon.shared.manifest_path}
61
+ index_path: results/indices/recon/recon_train.json
62
+ backend: ${recon.shared.backend}
63
+ sampler: ${recon.sampler.context_goal}
64
+ transform: ${recon.shared.transform}
65
+ test:
66
+ _target_: far.data.datasets.FormattedVideoDataset
67
+ formatter: ${recon.formatter.context_pose}
68
+ base_dataset:
69
+ _target_: far.data.datasets.VideoDataset
70
+ manifest_path: ${recon.shared.manifest_path}
71
+ index_path: results/indices/recon/recon_test.json
72
+ backend: ${recon.shared.backend}
73
+ sampler: ${recon.sampler.context_goal}
74
+ transform: ${recon.shared.transform}
75
+ eval:
76
+ _target_: far.data.datasets.FormattedVideoDataset
77
+ formatter: ${recon.formatter.context_pose}
78
+ base_dataset:
79
+ _target_: far.data.datasets.VideoDataset
80
+ manifest_path: ${recon.shared.manifest_path}
81
+ index_path: results/indices/recon/recon_test_time.json
82
+ backend: ${recon.shared.backend}
83
+ sampler: ${recon.sampler.context_traj}
84
+ transform: ${recon.shared.transform}
85
+ eval_traj:
86
+ _target_: far.data.datasets.FormattedVideoDataset
87
+ formatter: ${recon.formatter.context_pose}
88
+ base_dataset:
89
+ _target_: far.data.datasets.VideoDataset
90
+ manifest_path: ${recon.shared.manifest_path}
91
+ index_path: results/indices/recon/recon_test_plan.json
92
+ backend: ${recon.shared.backend}
93
+ sampler: ${recon.sampler.context_traj}
94
+ transform: ${recon.shared.transform}
95
+ loopnav:
96
+ spec:
97
+ _target_: far.builders.data.build_dataspec
98
+ local_pos_stats:
99
+ min:
100
+ - -2.5
101
+ - -4
102
+ - 0
103
+ max:
104
+ - 5
105
+ - 4
106
+ - 1
107
+ meters_per_waypoint: 0.25
108
+ max_frame_offset: 128
109
+ len_traj_pred: ${len_traj_pred}
110
+ planner_mu_init:
111
+ - -0.1
112
+ - 0
113
+ - 0
114
+ - 0
115
+ - 0
116
+ planner_sigma_init:
117
+ - 0.02
118
+ - 0.1
119
+ - 0
120
+ - 0
121
+ - 0.3142
122
+ pose_fields:
123
+ - x
124
+ - 'y'
125
+ - z
126
+ - pitch
127
+ - yaw
128
+ pose_dim: 5
129
+ pos_dim: 3
130
+ angle_dim: 2
131
+ pose_mode: 3d
132
+ fps: 20.0
133
+ shared:
134
+ manifest_path: results/manifests/loopnav/loopnav_split.json
135
+ backend:
136
+ _target_: far.data.backends.LoopNavBackend
137
+ standard_frame: false
138
+ transform:
139
+ _target_: far.data.transforms.make_resize
140
+ image_height: ${image_height}
141
+ image_width: ${image_width}
142
+ sampler:
143
+ random_clip:
144
+ _target_: far.data.samplers.RandomClipSampler
145
+ clip_len: 16
146
+ frame_stride: 10
147
+ context_goal:
148
+ _target_: far.data.samplers.ContextGoalSampler
149
+ context_pool_size: 100
150
+ len_traj_pred: ${len_traj_pred}
151
+ goals_per_obs: ${goals_per_obs}
152
+ context_traj:
153
+ _target_: far.data.samplers.ContextTrajectorySampler
154
+ context_pool_size: 100
155
+ len_traj_pred: ${len_traj_pred}
156
+ formatter:
157
+ frame:
158
+ _target_: far.data.formatters.FrameFormatter
159
+ context_pose:
160
+ _target_: far.data.formatters.ContextPoseFormatter
161
+ spec: ${loopnav.spec}
162
+ train_tokenizer:
163
+ _target_: far.data.datasets.FormattedIterableDataset
164
+ formatter: ${loopnav.formatter.frame}
165
+ base_dataset:
166
+ _target_: far.data.datasets.WebVideoIterableDataset
167
+ manifest_path: results/manifests/loopnav/loopnav.json
168
+ index_path: results/indices/loopnav/loopnav_train_ABCA.json
169
+ backend:
170
+ _target_: far.data.backends.WebDatasetBackend
171
+ sampler: ${loopnav.sampler.random_clip}
172
+ transform: ${loopnav.shared.transform}
173
+ cache_dir: null
174
+ cache_size: 100
175
+ shuffle_shards: 32
176
+ shuffle_samples: 100
177
+ loopnav_latent:
178
+ spec:
179
+ _target_: far.builders.data.build_dataspec
180
+ local_pos_stats:
181
+ min:
182
+ - 0
183
+ - 0
184
+ - 0
185
+ max:
186
+ - 1
187
+ - 1
188
+ - 1
189
+ meters_per_waypoint: 1.0
190
+ max_frame_offset: 20.0
191
+ len_traj_pred: ${len_traj_pred}
192
+ planner_mu_init:
193
+ - -0.1
194
+ - 0
195
+ - 0
196
+ - 0
197
+ - 0
198
+ planner_sigma_init:
199
+ - 0.02
200
+ - 0.1
201
+ - 0
202
+ - 0
203
+ - 0.3142
204
+ pose_fields:
205
+ - x
206
+ - 'y'
207
+ - z
208
+ - pitch
209
+ - yaw
210
+ pose_dim: 5
211
+ pos_dim: 3
212
+ angle_dim: 2
213
+ pose_mode: 3d
214
+ fps: 20.0
215
+ shared:
216
+ manifest_path: results/manifests/loopnav/loopnav_latent.json
217
+ backend:
218
+ _target_: far.data.backends.LatentLoopNavBackend
219
+ key_suffix: .keys.npy
220
+ standard_frame: false
221
+ transform: null
222
+ sampler:
223
+ random_clip:
224
+ _target_: far.data.samplers.RandomClipSampler
225
+ clip_len: 16
226
+ frame_stride: 10
227
+ context_goal:
228
+ _target_: far.data.samplers.ContextGoalSampler
229
+ context_pool_size: 200
230
+ len_traj_pred: ${len_traj_pred}
231
+ goals_per_obs: ${goals_per_obs}
232
+ context_goal_test:
233
+ _target_: far.data.samplers.ContextGoalSampler
234
+ context_pool_size: 200
235
+ len_traj_pred: ${len_traj_pred}
236
+ goals_per_obs: 4
237
+ context_traj:
238
+ _target_: far.data.samplers.ContextTrajectorySampler
239
+ context_pool_size: 200
240
+ len_traj_pred: ${len_traj_pred}
241
+ formatter:
242
+ frame:
243
+ _target_: far.data.formatters.FrameFormatter
244
+ context_pose:
245
+ _target_: far.data.formatters.ContextPoseFormatter
246
+ spec: ${loopnav_latent.spec}
247
+ train:
248
+ _target_: far.data.datasets.FormattedVideoDataset
249
+ formatter: ${loopnav_latent.formatter.context_pose}
250
+ base_dataset:
251
+ _target_: far.data.datasets.VideoDataset
252
+ manifest_path: ${loopnav_latent.shared.manifest_path}
253
+ index_path: results/indices/loopnav/loopnav_latent_train.json
254
+ backend: ${loopnav_latent.shared.backend}
255
+ sampler: ${loopnav_latent.sampler.context_goal}
256
+ transform: ${loopnav_latent.shared.transform}
257
+ test:
258
+ _target_: far.data.datasets.FormattedVideoDataset
259
+ formatter: ${loopnav_latent.formatter.context_pose}
260
+ base_dataset:
261
+ _target_: far.data.datasets.VideoDataset
262
+ manifest_path: ${loopnav_latent.shared.manifest_path}
263
+ index_path: results/indices/loopnav/loopnav_latent_test.json
264
+ backend: ${loopnav_latent.shared.backend}
265
+ sampler: ${loopnav_latent.sampler.context_goal_test}
266
+ transform: ${loopnav_latent.shared.transform}
267
+ soundspaces_latent:
268
+ spec:
269
+ _target_: far.builders.data.build_dataspec
270
+ local_pos_stats:
271
+ min:
272
+ - 0
273
+ - 0
274
+ - 0
275
+ max:
276
+ - 1
277
+ - 1
278
+ - 1
279
+ meters_per_waypoint: 1.0
280
+ max_frame_offset: 10.0
281
+ len_traj_pred: ${len_traj_pred}
282
+ planner_mu_init:
283
+ - -0.1
284
+ - 0
285
+ - 0
286
+ - 0
287
+ - 0
288
+ planner_sigma_init:
289
+ - 0.02
290
+ - 0.1
291
+ - 0
292
+ - 0
293
+ - 0.3142
294
+ pose_fields:
295
+ - x
296
+ - 'y'
297
+ - z
298
+ - pitch
299
+ - yaw
300
+ pose_dim: 5
301
+ pos_dim: 3
302
+ angle_dim: 2
303
+ pose_mode: 3d
304
+ fps: 10.0
305
+ shared:
306
+ manifest_path: results/manifests/soundspaces/soundspaces_latent.json
307
+ backend:
308
+ _target_: far.data.backends.SoundSpacesBackend
309
+ key_suffix: null
310
+ audio_mode: null
311
+ audio_window_hops: 48
312
+ transform: null
313
+ sampler:
314
+ random_clip:
315
+ _target_: far.data.samplers.RandomClipSampler
316
+ clip_len: 16
317
+ frame_stride: 5
318
+ context_goal:
319
+ _target_: far.data.samplers.ContextGoalSampler
320
+ context_pool_size: 200
321
+ len_traj_pred: ${len_traj_pred}
322
+ goals_per_obs: ${goals_per_obs}
323
+ context_goal_test:
324
+ _target_: far.data.samplers.ContextGoalSampler
325
+ context_pool_size: 200
326
+ len_traj_pred: ${len_traj_pred}
327
+ goals_per_obs: 4
328
+ context_traj:
329
+ _target_: far.data.samplers.ContextTrajectorySampler
330
+ context_pool_size: 200
331
+ len_traj_pred: ${len_traj_pred}
332
+ formatter:
333
+ frame:
334
+ _target_: far.data.formatters.FrameFormatter
335
+ context_pose:
336
+ _target_: far.data.formatters.ContextPoseFormatter
337
+ spec: ${soundspaces_latent.spec}
338
+ train:
339
+ _target_: far.data.datasets.FormattedVideoDataset
340
+ formatter: ${soundspaces_latent.formatter.context_pose}
341
+ base_dataset:
342
+ _target_: far.data.datasets.VideoDataset
343
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
344
+ index_path: results/indices/soundspaces/soundspaces_latent_train.json
345
+ backend: ${soundspaces_latent.shared.backend}
346
+ sampler: ${soundspaces_latent.sampler.context_goal}
347
+ transform: ${soundspaces_latent.shared.transform}
348
+ test:
349
+ _target_: far.data.datasets.FormattedVideoDataset
350
+ formatter: ${soundspaces_latent.formatter.context_pose}
351
+ base_dataset:
352
+ _target_: far.data.datasets.VideoDataset
353
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
354
+ index_path: results/indices/soundspaces/soundspaces_latent_test.json
355
+ backend: ${soundspaces_latent.shared.backend}
356
+ sampler: ${soundspaces_latent.sampler.context_goal_test}
357
+ transform: ${soundspaces_latent.shared.transform}
358
+ soundspaces_v2_latent:
359
+ spec:
360
+ _target_: far.builders.data.build_dataspec
361
+ local_pos_stats:
362
+ min:
363
+ - 0
364
+ - 0
365
+ - 0
366
+ max:
367
+ - 1
368
+ - 1
369
+ - 1
370
+ meters_per_waypoint: 1.0
371
+ max_frame_offset: 10.0
372
+ len_traj_pred: ${len_traj_pred}
373
+ planner_mu_init:
374
+ - -0.1
375
+ - 0
376
+ - 0
377
+ - 0
378
+ - 0
379
+ planner_sigma_init:
380
+ - 0.02
381
+ - 0.1
382
+ - 0
383
+ - 0
384
+ - 0.3142
385
+ pose_fields:
386
+ - x
387
+ - 'y'
388
+ - z
389
+ - pitch
390
+ - yaw
391
+ pose_dim: 5
392
+ pos_dim: 3
393
+ angle_dim: 2
394
+ pose_mode: 3d
395
+ fps: 10.0
396
+ shared:
397
+ manifest_path: results/manifests/soundspaces_v2/soundspaces_v2_latent.json
398
+ backend:
399
+ _target_: far.data.backends.SoundSpacesBackend
400
+ key_suffix: null
401
+ audio_mode: null
402
+ audio_window_hops: 48
403
+ transform: null
404
+ sampler:
405
+ random_clip:
406
+ _target_: far.data.samplers.RandomClipSampler
407
+ clip_len: 16
408
+ frame_stride: 5
409
+ context_goal:
410
+ _target_: far.data.samplers.ContextGoalSampler
411
+ context_pool_size: 200
412
+ len_traj_pred: ${len_traj_pred}
413
+ goals_per_obs: ${goals_per_obs}
414
+ context_goal_test:
415
+ _target_: far.data.samplers.ContextGoalSampler
416
+ context_pool_size: 200
417
+ len_traj_pred: ${len_traj_pred}
418
+ goals_per_obs: 4
419
+ context_traj:
420
+ _target_: far.data.samplers.ContextTrajectorySampler
421
+ context_pool_size: 200
422
+ len_traj_pred: ${len_traj_pred}
423
+ formatter:
424
+ frame:
425
+ _target_: far.data.formatters.FrameFormatter
426
+ context_pose:
427
+ _target_: far.data.formatters.ContextPoseFormatter
428
+ spec: ${soundspaces_v2_latent.spec}
429
+ train:
430
+ _target_: far.data.datasets.FormattedVideoDataset
431
+ formatter: ${soundspaces_v2_latent.formatter.context_pose}
432
+ base_dataset:
433
+ _target_: far.data.datasets.VideoDataset
434
+ manifest_path: ${soundspaces_v2_latent.shared.manifest_path}
435
+ index_path: results/indices/soundspaces_v2/soundspaces_v2_latent_train.json
436
+ backend: ${soundspaces_v2_latent.shared.backend}
437
+ sampler: ${soundspaces_v2_latent.sampler.context_goal}
438
+ transform: ${soundspaces_v2_latent.shared.transform}
439
+ test:
440
+ _target_: far.data.datasets.FormattedVideoDataset
441
+ formatter: ${soundspaces_v2_latent.formatter.context_pose}
442
+ base_dataset:
443
+ _target_: far.data.datasets.VideoDataset
444
+ manifest_path: ${soundspaces_v2_latent.shared.manifest_path}
445
+ index_path: results/indices/soundspaces_v2/soundspaces_v2_latent_test.json
446
+ backend: ${soundspaces_v2_latent.shared.backend}
447
+ sampler: ${soundspaces_v2_latent.sampler.context_goal_test}
448
+ transform: ${soundspaces_v2_latent.shared.transform}
449
+ ai2thor_latent:
450
+ spec:
451
+ _target_: far.builders.data.build_dataspec
452
+ local_pos_stats:
453
+ min:
454
+ - 0
455
+ - 0
456
+ - 0
457
+ max:
458
+ - 1
459
+ - 1
460
+ - 1
461
+ meters_per_waypoint: 1.0
462
+ max_frame_offset: 10.0
463
+ len_traj_pred: ${len_traj_pred}
464
+ planner_mu_init:
465
+ - -0.1
466
+ - 0
467
+ - 0
468
+ - 0
469
+ - 0
470
+ planner_sigma_init:
471
+ - 0.02
472
+ - 0.1
473
+ - 0
474
+ - 0
475
+ - 0.3142
476
+ pose_fields:
477
+ - x
478
+ - 'y'
479
+ - z
480
+ - pitch
481
+ - yaw
482
+ pose_dim: 5
483
+ pos_dim: 3
484
+ angle_dim: 2
485
+ pose_mode: 3d
486
+ fps: 10.0
487
+ shared:
488
+ manifest_path: results/manifests/ai2thor/ai2thor_latent.json
489
+ backend:
490
+ _target_: far.data.backends.THORBackend
491
+ key_suffix: null
492
+ transform: null
493
+ sampler:
494
+ random_clip:
495
+ _target_: far.data.samplers.RandomClipSampler
496
+ clip_len: 16
497
+ frame_stride: 5
498
+ context_goal:
499
+ _target_: far.data.samplers.ContextGoalSampler
500
+ context_pool_size: 200
501
+ len_traj_pred: ${len_traj_pred}
502
+ goals_per_obs: ${goals_per_obs}
503
+ context_goal_test:
504
+ _target_: far.data.samplers.ContextGoalSampler
505
+ context_pool_size: 200
506
+ len_traj_pred: ${len_traj_pred}
507
+ goals_per_obs: 4
508
+ context_traj:
509
+ _target_: far.data.samplers.ContextTrajectorySampler
510
+ context_pool_size: 200
511
+ len_traj_pred: ${len_traj_pred}
512
+ formatter:
513
+ frame:
514
+ _target_: far.data.formatters.FrameFormatter
515
+ context_pose:
516
+ _target_: far.data.formatters.ContextPoseFormatter
517
+ spec: ${ai2thor_latent.spec}
518
+ train:
519
+ _target_: far.data.datasets.FormattedVideoDataset
520
+ formatter: ${ai2thor_latent.formatter.context_pose}
521
+ base_dataset:
522
+ _target_: far.data.datasets.VideoDataset
523
+ manifest_path: ${ai2thor_latent.shared.manifest_path}
524
+ index_path: results/indices/ai2thor/ai2thor_latent_train.json
525
+ backend: ${ai2thor_latent.shared.backend}
526
+ sampler: ${ai2thor_latent.sampler.context_goal}
527
+ transform: ${ai2thor_latent.shared.transform}
528
+ test:
529
+ _target_: far.data.datasets.FormattedVideoDataset
530
+ formatter: ${ai2thor_latent.formatter.context_pose}
531
+ base_dataset:
532
+ _target_: far.data.datasets.VideoDataset
533
+ manifest_path: ${ai2thor_latent.shared.manifest_path}
534
+ index_path: results/indices/ai2thor/ai2thor_latent_test.json
535
+ backend: ${ai2thor_latent.shared.backend}
536
+ sampler: ${ai2thor_latent.sampler.context_goal_test}
537
+ transform: ${ai2thor_latent.shared.transform}
538
+ ai2thor_v2_latent:
539
+ spec:
540
+ _target_: far.builders.data.build_dataspec
541
+ local_pos_stats:
542
+ min:
543
+ - 0
544
+ - 0
545
+ - 0
546
+ max:
547
+ - 1
548
+ - 1
549
+ - 1
550
+ meters_per_waypoint: 1.0
551
+ max_frame_offset: 10.0
552
+ len_traj_pred: ${len_traj_pred}
553
+ planner_mu_init:
554
+ - -0.1
555
+ - 0
556
+ - 0
557
+ - 0
558
+ - 0
559
+ planner_sigma_init:
560
+ - 0.02
561
+ - 0.1
562
+ - 0
563
+ - 0
564
+ - 0.3142
565
+ pose_fields:
566
+ - x
567
+ - 'y'
568
+ - z
569
+ - pitch
570
+ - yaw
571
+ pose_dim: 5
572
+ pos_dim: 3
573
+ angle_dim: 2
574
+ pose_mode: 3d
575
+ fps: 10.0
576
+ shared:
577
+ manifest_path: results/manifests/ai2thor_v2/ai2thor_v2_latent.json
578
+ backend:
579
+ _target_: far.data.backends.THORBackend
580
+ key_suffix: null
581
+ mask_suffix: null
582
+ transform: null
583
+ sampler:
584
+ random_clip:
585
+ _target_: far.data.samplers.RandomClipSampler
586
+ clip_len: 16
587
+ frame_stride: 5
588
+ context_goal:
589
+ _target_: far.data.samplers.ContextGoalSampler
590
+ context_pool_size: 200
591
+ len_traj_pred: ${len_traj_pred}
592
+ goals_per_obs: ${goals_per_obs}
593
+ event_bias_p: 0.0
594
+ event_phase: eval
595
+ event_codes: null
596
+ event_pre:
597
+ - -3
598
+ - 5
599
+ event_post:
600
+ - 0
601
+ - 8
602
+ event_goal_anchor: start
603
+ pool_phase: pre_eval
604
+ context_goal_test:
605
+ _target_: far.data.samplers.ContextGoalSampler
606
+ context_pool_size: 200
607
+ len_traj_pred: ${len_traj_pred}
608
+ goals_per_obs: 4
609
+ pool_phase: pre_eval
610
+ context_traj:
611
+ _target_: far.data.samplers.ContextTrajectorySampler
612
+ context_pool_size: 200
613
+ len_traj_pred: ${len_traj_pred}
614
+ pool_phase: pre_eval
615
+ formatter:
616
+ frame:
617
+ _target_: far.data.formatters.FrameFormatter
618
+ context_pose:
619
+ _target_: far.data.formatters.ContextPoseFormatter
620
+ spec: ${ai2thor_v2_latent.spec}
621
+ train:
622
+ _target_: far.data.datasets.FormattedVideoDataset
623
+ formatter: ${ai2thor_v2_latent.formatter.context_pose}
624
+ base_dataset:
625
+ _target_: far.data.datasets.VideoDataset
626
+ manifest_path: ${ai2thor_v2_latent.shared.manifest_path}
627
+ index_path: results/indices/ai2thor_v2/ai2thor_v2_latent_train.json
628
+ backend: ${ai2thor_v2_latent.shared.backend}
629
+ sampler: ${ai2thor_v2_latent.sampler.context_goal}
630
+ transform: ${ai2thor_v2_latent.shared.transform}
631
+ event_index_path: results/indices/ai2thor_v2/ai2thor_v2_latent_train_events.json
632
+ event_frac: null
633
+ test:
634
+ _target_: far.data.datasets.FormattedVideoDataset
635
+ formatter: ${ai2thor_v2_latent.formatter.context_pose}
636
+ base_dataset:
637
+ _target_: far.data.datasets.VideoDataset
638
+ manifest_path: ${ai2thor_v2_latent.shared.manifest_path}
639
+ index_path: results/indices/ai2thor_v2/ai2thor_v2_latent_test.json
640
+ backend: ${ai2thor_v2_latent.shared.backend}
641
+ sampler: ${ai2thor_v2_latent.sampler.context_goal_test}
642
+ transform: ${ai2thor_v2_latent.shared.transform}
643
+ ai2thor_v3_latent:
644
+ spec:
645
+ _target_: far.builders.data.build_dataspec
646
+ local_pos_stats:
647
+ min:
648
+ - 0
649
+ - 0
650
+ - 0
651
+ max:
652
+ - 1
653
+ - 1
654
+ - 1
655
+ meters_per_waypoint: 1.0
656
+ max_frame_offset: 10.0
657
+ len_traj_pred: ${len_traj_pred}
658
+ planner_mu_init:
659
+ - -0.1
660
+ - 0
661
+ - 0
662
+ - 0
663
+ - 0
664
+ planner_sigma_init:
665
+ - 0.02
666
+ - 0.1
667
+ - 0
668
+ - 0
669
+ - 0.3142
670
+ pose_fields:
671
+ - x
672
+ - 'y'
673
+ - z
674
+ - pitch
675
+ - yaw
676
+ pose_dim: 5
677
+ pos_dim: 3
678
+ angle_dim: 2
679
+ pose_mode: 3d
680
+ fps: 10.0
681
+ shared:
682
+ manifest_path: results/manifests/ai2thor_v3/ai2thor_v3_latent.json
683
+ backend:
684
+ _target_: far.data.backends.THORBackend
685
+ key_suffix: null
686
+ mask_suffix: null
687
+ transform: null
688
+ sampler:
689
+ random_clip:
690
+ _target_: far.data.samplers.RandomClipSampler
691
+ clip_len: 16
692
+ frame_stride: 5
693
+ context_goal:
694
+ _target_: far.data.samplers.ContextGoalSampler
695
+ context_pool_size: 200
696
+ len_traj_pred: ${len_traj_pred}
697
+ goals_per_obs: ${goals_per_obs}
698
+ event_bias_p: 0.0
699
+ event_phase: eval
700
+ event_codes: null
701
+ event_pre:
702
+ - -3
703
+ - 5
704
+ event_post:
705
+ - 0
706
+ - 8
707
+ event_goal_anchor: start
708
+ pool_phase: pre_eval
709
+ context_goal_test:
710
+ _target_: far.data.samplers.ContextGoalSampler
711
+ context_pool_size: 200
712
+ len_traj_pred: ${len_traj_pred}
713
+ goals_per_obs: 4
714
+ pool_phase: pre_eval
715
+ context_traj:
716
+ _target_: far.data.samplers.ContextTrajectorySampler
717
+ context_pool_size: 200
718
+ len_traj_pred: ${len_traj_pred}
719
+ pool_phase: pre_eval
720
+ formatter:
721
+ frame:
722
+ _target_: far.data.formatters.FrameFormatter
723
+ context_pose:
724
+ _target_: far.data.formatters.ContextPoseFormatter
725
+ spec: ${ai2thor_v3_latent.spec}
726
+ train:
727
+ _target_: far.data.datasets.FormattedVideoDataset
728
+ formatter: ${ai2thor_v3_latent.formatter.context_pose}
729
+ base_dataset:
730
+ _target_: far.data.datasets.VideoDataset
731
+ manifest_path: ${ai2thor_v3_latent.shared.manifest_path}
732
+ index_path: results/indices/ai2thor_v3/ai2thor_v3_latent_train.json
733
+ backend: ${ai2thor_v3_latent.shared.backend}
734
+ sampler: ${ai2thor_v3_latent.sampler.context_goal}
735
+ transform: ${ai2thor_v3_latent.shared.transform}
736
+ event_index_path: results/indices/ai2thor_v3/ai2thor_v3_latent_train_events.json
737
+ event_frac: null
738
+ test:
739
+ _target_: far.data.datasets.FormattedVideoDataset
740
+ formatter: ${ai2thor_v3_latent.formatter.context_pose}
741
+ base_dataset:
742
+ _target_: far.data.datasets.VideoDataset
743
+ manifest_path: ${ai2thor_v3_latent.shared.manifest_path}
744
+ index_path: results/indices/ai2thor_v3/ai2thor_v3_latent_test.json
745
+ backend: ${ai2thor_v3_latent.shared.backend}
746
+ sampler: ${ai2thor_v3_latent.sampler.context_goal_test}
747
+ transform: ${ai2thor_v3_latent.shared.transform}
748
+ ai2thor_dyn_latent:
749
+ spec:
750
+ _target_: far.builders.data.build_dataspec
751
+ local_pos_stats:
752
+ min:
753
+ - 0
754
+ - 0
755
+ - 0
756
+ max:
757
+ - 1
758
+ - 1
759
+ - 1
760
+ meters_per_waypoint: 1.0
761
+ max_frame_offset: 10.0
762
+ len_traj_pred: ${len_traj_pred}
763
+ planner_mu_init:
764
+ - -0.1
765
+ - 0
766
+ - 0
767
+ - 0
768
+ - 0
769
+ planner_sigma_init:
770
+ - 0.02
771
+ - 0.1
772
+ - 0
773
+ - 0
774
+ - 0.3142
775
+ pose_fields:
776
+ - x
777
+ - 'y'
778
+ - z
779
+ - pitch
780
+ - yaw
781
+ pose_dim: 5
782
+ pos_dim: 3
783
+ angle_dim: 2
784
+ pose_mode: 3d
785
+ fps: 10.0
786
+ shared:
787
+ manifest_path: results/manifests/ai2thor_dyn/ai2thor_dyn_latent.json
788
+ backend:
789
+ _target_: far.data.backends.THORBackend
790
+ key_suffix: null
791
+ mask_suffix: .agent.npy
792
+ transform: null
793
+ sampler:
794
+ random_clip:
795
+ _target_: far.data.samplers.RandomClipSampler
796
+ clip_len: 16
797
+ frame_stride: 5
798
+ context_goal:
799
+ _target_: far.data.samplers.ContextGoalSampler
800
+ context_pool_size: 400
801
+ len_traj_pred: ${len_traj_pred}
802
+ goals_per_obs: ${goals_per_obs}
803
+ event_bias_p: 0.0
804
+ event_phase: eval
805
+ event_codes: null
806
+ event_pre:
807
+ - -3
808
+ - 5
809
+ event_post:
810
+ - 0
811
+ - 8
812
+ event_goal_anchor: start
813
+ pool_phase: pre_eval
814
+ long_horizon_p: 0.0
815
+ long_horizon_min_past: 50
816
+ context_goal_test:
817
+ _target_: far.data.samplers.ContextGoalSampler
818
+ context_pool_size: 400
819
+ len_traj_pred: ${len_traj_pred}
820
+ goals_per_obs: 4
821
+ pool_phase: pre_eval
822
+ long_horizon_p: 0.0
823
+ long_horizon_min_past: 50
824
+ context_traj:
825
+ _target_: far.data.samplers.ContextTrajectorySampler
826
+ context_pool_size: 400
827
+ len_traj_pred: ${len_traj_pred}
828
+ pool_phase: pre_eval
829
+ formatter:
830
+ frame:
831
+ _target_: far.data.formatters.FrameFormatter
832
+ context_pose:
833
+ _target_: far.data.formatters.ContextPoseFormatter
834
+ spec: ${ai2thor_dyn_latent.spec}
835
+ train:
836
+ _target_: far.data.datasets.FormattedVideoDataset
837
+ formatter: ${ai2thor_dyn_latent.formatter.context_pose}
838
+ base_dataset:
839
+ _target_: far.data.datasets.VideoDataset
840
+ manifest_path: ${ai2thor_dyn_latent.shared.manifest_path}
841
+ index_path: results/indices/ai2thor_dyn/ai2thor_dyn_latent_train.json
842
+ backend: ${ai2thor_dyn_latent.shared.backend}
843
+ sampler: ${ai2thor_dyn_latent.sampler.context_goal}
844
+ transform: ${ai2thor_dyn_latent.shared.transform}
845
+ event_index_path: null
846
+ event_frac: null
847
+ test:
848
+ _target_: far.data.datasets.FormattedVideoDataset
849
+ formatter: ${ai2thor_dyn_latent.formatter.context_pose}
850
+ base_dataset:
851
+ _target_: far.data.datasets.VideoDataset
852
+ manifest_path: ${ai2thor_dyn_latent.shared.manifest_path}
853
+ index_path: results/indices/ai2thor_dyn/ai2thor_dyn_latent_test.json
854
+ backend: ${ai2thor_dyn_latent.shared.backend}
855
+ sampler: ${ai2thor_dyn_latent.sampler.context_goal_test}
856
+ transform: ${ai2thor_dyn_latent.shared.transform}
857
+ model:
858
+ memory:
859
+ _target_: far.memories.pnp_cl.PnPRetrieval
860
+ ckpt_path: null
861
+ chunk_size: 10
862
+ num_sets: 4
863
+ softmax_temperature: 1.0
864
+ scoring: visual
865
+ metadata_hidden: 128
866
+ metadata_layers: 2
867
+ local_pos_stats: null
868
+ stride_to_seconds: 1.0
869
+ encoder_mode: frozen
870
+ adapter_hidden: 384
871
+ query_source: frame
872
+ query_conditioning: action
873
+ gate_eps: 0.05
874
+ query_vit:
875
+ _target_: far.memories.pnp_cl.QueryViT
876
+ img_height: 32
877
+ img_width: 32
878
+ patch_size: 2
879
+ in_chans: 4
880
+ embed_dim: 384
881
+ depth: 6
882
+ num_heads: 12
883
+ action_dim: ${action_dim}
884
+ key_dim: 256
885
+ mlp_ratio: 4.0
886
+ interact_dim: 4
887
+ wandb:
888
+ project: PMem-THOR
889
+ name: null
890
+ entity: null
891
+ mode: online
892
+ group: ai2thor_dyn_latent
893
+ tags:
894
+ - ai2thor_dyn_latent
895
+ - retriever_pretrain
896
+ - object_centric
897
+ run_name: retriever_thordyn_object_ft
898
+ context_size: 4
899
+ action_dim: 5
900
+ pose_dim: 5
901
+ image_height: 256
902
+ image_width: 256
903
+ len_traj_pred: 64
904
+ goals_per_obs: 1
905
+ lr: 5.0e-05
906
+ weight_decay: 0.0
907
+ grad_clip_val: 1.0
908
+ iterations: 50000
909
+ warmup_steps: 2000
910
+ from_checkpoint: null
911
+ global_seed: 0
912
+ log_every: 100
913
+ ckpt_every: 2500
914
+ eval_every: 2500
915
+ bfloat16: false
916
+ batch_size: 16
917
+ num_workers: 12
918
+ datasets:
919
+ - ai2thor_dyn_latent
920
+ contrastive:
921
+ chunk_size: ${model.memory.chunk_size}
922
+ positive_pair_type:
923
+ - jepa
924
+ num_positives: 8
925
+ num_negatives: 24
926
+ cl_temperature: 0.1
927
+ cl_pose_radius: 0.5
928
+ cl_pose_fov_half_h: 180.0
929
+ cl_pose_fov_half_v: 180.0
930
+ cl_pose_num_samples: 1024
931
+ local_pos_stats: ${ai2thor_dyn_latent.spec.local_pos_stats}
932
+ init_from: results/20260818_150341__pretrain_retriever__retriever_thor_jepa_pretrain_ltp64__bs16__ai2thor_latent/checkpoints/0400000.pth.tar
933
+ obj_weight: 1.0
934
+ obj_num_positives: 8
935
+ obj_pulse_gap_s: 2.0
936
+ obj_min_frac: 0.0
937
+ obj_last_only: true
938
+ num_hard_negatives: 0
939
+ hard_neg_weight: 4.0
940
+ hard_pos_tol: 0.05
941
+ hard_yaw_tol_deg: 1.0
942
+ hard_pitch_tol_deg: 1.0
943
+ hard_tau: 0.05
944
+ hard_min_tokens: 8
945
+ hard_sym_weight: 1.0
946
+ stride_clip: null
947
+ max_frame_offset: 10.0
948
+ fps: 10.0
ai2thor_v3/retriever_jepa/checkpoints/0400000.pth.tar ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:32ecb108654a7df2b3404c9508f4ada791c8c802b48558ab978eae5d9830119d
3
+ size 66429235
ai2thor_v3/retriever_jepa/config.yaml ADDED
@@ -0,0 +1,611 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ recon:
2
+ spec:
3
+ _target_: far.builders.data.build_dataspec
4
+ local_pos_stats:
5
+ min:
6
+ - -2.5
7
+ - -4
8
+ max:
9
+ - 5
10
+ - 4
11
+ meters_per_waypoint: 0.25
12
+ max_frame_offset: 128
13
+ len_traj_pred: ${len_traj_pred}
14
+ planner_mu_init:
15
+ - -0.1
16
+ - 0
17
+ - 0
18
+ planner_sigma_init:
19
+ - 0.02
20
+ - 0.1
21
+ - 0.3142
22
+ pose_fields:
23
+ - x
24
+ - 'y'
25
+ - z
26
+ - pitch
27
+ - yaw
28
+ pose_dim: 5
29
+ pos_dim: 2
30
+ angle_dim: 1
31
+ pose_mode: 2d
32
+ fps: 4.0
33
+ shared:
34
+ manifest_path: results/manifests/recon/recon_split.json
35
+ backend:
36
+ _target_: far.data.backends.NavigationBackend
37
+ transform:
38
+ _target_: far.data.transforms.make_centercrop_resize
39
+ image_height: ${image_height}
40
+ image_width: ${image_width}
41
+ sampler:
42
+ context_goal:
43
+ _target_: far.data.samplers.ContextGoalSampler
44
+ context_pool_size: 12
45
+ len_traj_pred: ${len_traj_pred}
46
+ goals_per_obs: ${goals_per_obs}
47
+ context_traj:
48
+ _target_: far.data.samplers.ContextTrajectorySampler
49
+ context_pool_size: 12
50
+ len_traj_pred: ${len_traj_pred}
51
+ formatter:
52
+ context_pose:
53
+ _target_: far.data.formatters.ContextPoseFormatter2D
54
+ spec: ${recon.spec}
55
+ train:
56
+ _target_: far.data.datasets.FormattedVideoDataset
57
+ formatter: ${recon.formatter.context_pose}
58
+ base_dataset:
59
+ _target_: far.data.datasets.VideoDataset
60
+ manifest_path: ${recon.shared.manifest_path}
61
+ index_path: results/indices/recon/recon_train.json
62
+ backend: ${recon.shared.backend}
63
+ sampler: ${recon.sampler.context_goal}
64
+ transform: ${recon.shared.transform}
65
+ test:
66
+ _target_: far.data.datasets.FormattedVideoDataset
67
+ formatter: ${recon.formatter.context_pose}
68
+ base_dataset:
69
+ _target_: far.data.datasets.VideoDataset
70
+ manifest_path: ${recon.shared.manifest_path}
71
+ index_path: results/indices/recon/recon_test.json
72
+ backend: ${recon.shared.backend}
73
+ sampler: ${recon.sampler.context_goal}
74
+ transform: ${recon.shared.transform}
75
+ eval:
76
+ _target_: far.data.datasets.FormattedVideoDataset
77
+ formatter: ${recon.formatter.context_pose}
78
+ base_dataset:
79
+ _target_: far.data.datasets.VideoDataset
80
+ manifest_path: ${recon.shared.manifest_path}
81
+ index_path: results/indices/recon/recon_test_time.json
82
+ backend: ${recon.shared.backend}
83
+ sampler: ${recon.sampler.context_traj}
84
+ transform: ${recon.shared.transform}
85
+ eval_traj:
86
+ _target_: far.data.datasets.FormattedVideoDataset
87
+ formatter: ${recon.formatter.context_pose}
88
+ base_dataset:
89
+ _target_: far.data.datasets.VideoDataset
90
+ manifest_path: ${recon.shared.manifest_path}
91
+ index_path: results/indices/recon/recon_test_plan.json
92
+ backend: ${recon.shared.backend}
93
+ sampler: ${recon.sampler.context_traj}
94
+ transform: ${recon.shared.transform}
95
+ loopnav:
96
+ spec:
97
+ _target_: far.builders.data.build_dataspec
98
+ local_pos_stats:
99
+ min:
100
+ - -2.5
101
+ - -4
102
+ - 0
103
+ max:
104
+ - 5
105
+ - 4
106
+ - 1
107
+ meters_per_waypoint: 0.25
108
+ max_frame_offset: 128
109
+ len_traj_pred: ${len_traj_pred}
110
+ planner_mu_init:
111
+ - -0.1
112
+ - 0
113
+ - 0
114
+ - 0
115
+ - 0
116
+ planner_sigma_init:
117
+ - 0.02
118
+ - 0.1
119
+ - 0
120
+ - 0
121
+ - 0.3142
122
+ pose_fields:
123
+ - x
124
+ - 'y'
125
+ - z
126
+ - pitch
127
+ - yaw
128
+ pose_dim: 5
129
+ pos_dim: 3
130
+ angle_dim: 2
131
+ pose_mode: 3d
132
+ fps: 20.0
133
+ shared:
134
+ manifest_path: results/manifests/loopnav/loopnav_split.json
135
+ backend:
136
+ _target_: far.data.backends.LoopNavBackend
137
+ standard_frame: false
138
+ transform:
139
+ _target_: far.data.transforms.make_resize
140
+ image_height: ${image_height}
141
+ image_width: ${image_width}
142
+ sampler:
143
+ random_clip:
144
+ _target_: far.data.samplers.RandomClipSampler
145
+ clip_len: 16
146
+ frame_stride: 10
147
+ context_goal:
148
+ _target_: far.data.samplers.ContextGoalSampler
149
+ context_pool_size: 100
150
+ len_traj_pred: ${len_traj_pred}
151
+ goals_per_obs: ${goals_per_obs}
152
+ context_traj:
153
+ _target_: far.data.samplers.ContextTrajectorySampler
154
+ context_pool_size: 100
155
+ len_traj_pred: ${len_traj_pred}
156
+ formatter:
157
+ frame:
158
+ _target_: far.data.formatters.FrameFormatter
159
+ context_pose:
160
+ _target_: far.data.formatters.ContextPoseFormatter
161
+ spec: ${loopnav.spec}
162
+ train_tokenizer:
163
+ _target_: far.data.datasets.FormattedIterableDataset
164
+ formatter: ${loopnav.formatter.frame}
165
+ base_dataset:
166
+ _target_: far.data.datasets.WebVideoIterableDataset
167
+ manifest_path: results/manifests/loopnav/loopnav.json
168
+ index_path: results/indices/loopnav/loopnav_train_ABCA.json
169
+ backend:
170
+ _target_: far.data.backends.WebDatasetBackend
171
+ sampler: ${loopnav.sampler.random_clip}
172
+ transform: ${loopnav.shared.transform}
173
+ cache_dir: null
174
+ cache_size: 100
175
+ shuffle_shards: 32
176
+ shuffle_samples: 100
177
+ loopnav_latent:
178
+ spec:
179
+ _target_: far.builders.data.build_dataspec
180
+ local_pos_stats:
181
+ min:
182
+ - 0
183
+ - 0
184
+ - 0
185
+ max:
186
+ - 1
187
+ - 1
188
+ - 1
189
+ meters_per_waypoint: 1.0
190
+ max_frame_offset: 20.0
191
+ len_traj_pred: ${len_traj_pred}
192
+ planner_mu_init:
193
+ - -0.1
194
+ - 0
195
+ - 0
196
+ - 0
197
+ - 0
198
+ planner_sigma_init:
199
+ - 0.02
200
+ - 0.1
201
+ - 0
202
+ - 0
203
+ - 0.3142
204
+ pose_fields:
205
+ - x
206
+ - 'y'
207
+ - z
208
+ - pitch
209
+ - yaw
210
+ pose_dim: 5
211
+ pos_dim: 3
212
+ angle_dim: 2
213
+ pose_mode: 3d
214
+ fps: 20.0
215
+ shared:
216
+ manifest_path: results/manifests/loopnav/loopnav_latent.json
217
+ backend:
218
+ _target_: far.data.backends.LatentLoopNavBackend
219
+ key_suffix: .keys.npy
220
+ standard_frame: false
221
+ transform: null
222
+ sampler:
223
+ random_clip:
224
+ _target_: far.data.samplers.RandomClipSampler
225
+ clip_len: 16
226
+ frame_stride: 10
227
+ context_goal:
228
+ _target_: far.data.samplers.ContextGoalSampler
229
+ context_pool_size: 200
230
+ len_traj_pred: ${len_traj_pred}
231
+ goals_per_obs: ${goals_per_obs}
232
+ context_goal_test:
233
+ _target_: far.data.samplers.ContextGoalSampler
234
+ context_pool_size: 200
235
+ len_traj_pred: ${len_traj_pred}
236
+ goals_per_obs: 4
237
+ context_traj:
238
+ _target_: far.data.samplers.ContextTrajectorySampler
239
+ context_pool_size: 200
240
+ len_traj_pred: ${len_traj_pred}
241
+ formatter:
242
+ frame:
243
+ _target_: far.data.formatters.FrameFormatter
244
+ context_pose:
245
+ _target_: far.data.formatters.ContextPoseFormatter
246
+ spec: ${loopnav_latent.spec}
247
+ train:
248
+ _target_: far.data.datasets.FormattedVideoDataset
249
+ formatter: ${loopnav_latent.formatter.context_pose}
250
+ base_dataset:
251
+ _target_: far.data.datasets.VideoDataset
252
+ manifest_path: ${loopnav_latent.shared.manifest_path}
253
+ index_path: results/indices/loopnav/loopnav_latent_train.json
254
+ backend: ${loopnav_latent.shared.backend}
255
+ sampler: ${loopnav_latent.sampler.context_goal}
256
+ transform: ${loopnav_latent.shared.transform}
257
+ test:
258
+ _target_: far.data.datasets.FormattedVideoDataset
259
+ formatter: ${loopnav_latent.formatter.context_pose}
260
+ base_dataset:
261
+ _target_: far.data.datasets.VideoDataset
262
+ manifest_path: ${loopnav_latent.shared.manifest_path}
263
+ index_path: results/indices/loopnav/loopnav_latent_test.json
264
+ backend: ${loopnav_latent.shared.backend}
265
+ sampler: ${loopnav_latent.sampler.context_goal_test}
266
+ transform: ${loopnav_latent.shared.transform}
267
+ soundspaces_latent:
268
+ spec:
269
+ _target_: far.builders.data.build_dataspec
270
+ local_pos_stats:
271
+ min:
272
+ - 0
273
+ - 0
274
+ - 0
275
+ max:
276
+ - 1
277
+ - 1
278
+ - 1
279
+ meters_per_waypoint: 1.0
280
+ max_frame_offset: 10.0
281
+ len_traj_pred: ${len_traj_pred}
282
+ planner_mu_init:
283
+ - -0.1
284
+ - 0
285
+ - 0
286
+ - 0
287
+ - 0
288
+ planner_sigma_init:
289
+ - 0.02
290
+ - 0.1
291
+ - 0
292
+ - 0
293
+ - 0.3142
294
+ pose_fields:
295
+ - x
296
+ - 'y'
297
+ - z
298
+ - pitch
299
+ - yaw
300
+ pose_dim: 5
301
+ pos_dim: 3
302
+ angle_dim: 2
303
+ pose_mode: 3d
304
+ fps: 10.0
305
+ shared:
306
+ manifest_path: results/manifests/soundspaces/soundspaces_latent.json
307
+ backend:
308
+ _target_: far.data.backends.SoundSpacesBackend
309
+ key_suffix: null
310
+ audio_mode: null
311
+ audio_window_hops: 48
312
+ transform: null
313
+ sampler:
314
+ random_clip:
315
+ _target_: far.data.samplers.RandomClipSampler
316
+ clip_len: 16
317
+ frame_stride: 5
318
+ context_goal:
319
+ _target_: far.data.samplers.ContextGoalSampler
320
+ context_pool_size: 200
321
+ len_traj_pred: ${len_traj_pred}
322
+ goals_per_obs: ${goals_per_obs}
323
+ context_goal_test:
324
+ _target_: far.data.samplers.ContextGoalSampler
325
+ context_pool_size: 200
326
+ len_traj_pred: ${len_traj_pred}
327
+ goals_per_obs: 4
328
+ context_traj:
329
+ _target_: far.data.samplers.ContextTrajectorySampler
330
+ context_pool_size: 200
331
+ len_traj_pred: ${len_traj_pred}
332
+ formatter:
333
+ frame:
334
+ _target_: far.data.formatters.FrameFormatter
335
+ context_pose:
336
+ _target_: far.data.formatters.ContextPoseFormatter
337
+ spec: ${soundspaces_latent.spec}
338
+ train:
339
+ _target_: far.data.datasets.FormattedVideoDataset
340
+ formatter: ${soundspaces_latent.formatter.context_pose}
341
+ base_dataset:
342
+ _target_: far.data.datasets.VideoDataset
343
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
344
+ index_path: results/indices/soundspaces/soundspaces_latent_train.json
345
+ backend: ${soundspaces_latent.shared.backend}
346
+ sampler: ${soundspaces_latent.sampler.context_goal}
347
+ transform: ${soundspaces_latent.shared.transform}
348
+ test:
349
+ _target_: far.data.datasets.FormattedVideoDataset
350
+ formatter: ${soundspaces_latent.formatter.context_pose}
351
+ base_dataset:
352
+ _target_: far.data.datasets.VideoDataset
353
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
354
+ index_path: results/indices/soundspaces/soundspaces_latent_test.json
355
+ backend: ${soundspaces_latent.shared.backend}
356
+ sampler: ${soundspaces_latent.sampler.context_goal_test}
357
+ transform: ${soundspaces_latent.shared.transform}
358
+ soundspaces_v2_latent:
359
+ spec:
360
+ _target_: far.builders.data.build_dataspec
361
+ local_pos_stats:
362
+ min:
363
+ - 0
364
+ - 0
365
+ - 0
366
+ max:
367
+ - 1
368
+ - 1
369
+ - 1
370
+ meters_per_waypoint: 1.0
371
+ max_frame_offset: 10.0
372
+ len_traj_pred: ${len_traj_pred}
373
+ planner_mu_init:
374
+ - -0.1
375
+ - 0
376
+ - 0
377
+ - 0
378
+ - 0
379
+ planner_sigma_init:
380
+ - 0.02
381
+ - 0.1
382
+ - 0
383
+ - 0
384
+ - 0.3142
385
+ pose_fields:
386
+ - x
387
+ - 'y'
388
+ - z
389
+ - pitch
390
+ - yaw
391
+ pose_dim: 5
392
+ pos_dim: 3
393
+ angle_dim: 2
394
+ pose_mode: 3d
395
+ fps: 10.0
396
+ shared:
397
+ manifest_path: results/manifests/soundspaces_v2/soundspaces_v2_latent.json
398
+ backend:
399
+ _target_: far.data.backends.SoundSpacesBackend
400
+ key_suffix: null
401
+ audio_mode: null
402
+ audio_window_hops: 48
403
+ transform: null
404
+ sampler:
405
+ random_clip:
406
+ _target_: far.data.samplers.RandomClipSampler
407
+ clip_len: 16
408
+ frame_stride: 5
409
+ context_goal:
410
+ _target_: far.data.samplers.ContextGoalSampler
411
+ context_pool_size: 200
412
+ len_traj_pred: ${len_traj_pred}
413
+ goals_per_obs: ${goals_per_obs}
414
+ context_goal_test:
415
+ _target_: far.data.samplers.ContextGoalSampler
416
+ context_pool_size: 200
417
+ len_traj_pred: ${len_traj_pred}
418
+ goals_per_obs: 4
419
+ context_traj:
420
+ _target_: far.data.samplers.ContextTrajectorySampler
421
+ context_pool_size: 200
422
+ len_traj_pred: ${len_traj_pred}
423
+ formatter:
424
+ frame:
425
+ _target_: far.data.formatters.FrameFormatter
426
+ context_pose:
427
+ _target_: far.data.formatters.ContextPoseFormatter
428
+ spec: ${soundspaces_v2_latent.spec}
429
+ train:
430
+ _target_: far.data.datasets.FormattedVideoDataset
431
+ formatter: ${soundspaces_v2_latent.formatter.context_pose}
432
+ base_dataset:
433
+ _target_: far.data.datasets.VideoDataset
434
+ manifest_path: ${soundspaces_v2_latent.shared.manifest_path}
435
+ index_path: results/indices/soundspaces_v2/soundspaces_v2_latent_train.json
436
+ backend: ${soundspaces_v2_latent.shared.backend}
437
+ sampler: ${soundspaces_v2_latent.sampler.context_goal}
438
+ transform: ${soundspaces_v2_latent.shared.transform}
439
+ test:
440
+ _target_: far.data.datasets.FormattedVideoDataset
441
+ formatter: ${soundspaces_v2_latent.formatter.context_pose}
442
+ base_dataset:
443
+ _target_: far.data.datasets.VideoDataset
444
+ manifest_path: ${soundspaces_v2_latent.shared.manifest_path}
445
+ index_path: results/indices/soundspaces_v2/soundspaces_v2_latent_test.json
446
+ backend: ${soundspaces_v2_latent.shared.backend}
447
+ sampler: ${soundspaces_v2_latent.sampler.context_goal_test}
448
+ transform: ${soundspaces_v2_latent.shared.transform}
449
+ ai2thor_latent:
450
+ spec:
451
+ _target_: far.builders.data.build_dataspec
452
+ local_pos_stats:
453
+ min:
454
+ - 0
455
+ - 0
456
+ - 0
457
+ max:
458
+ - 1
459
+ - 1
460
+ - 1
461
+ meters_per_waypoint: 1.0
462
+ max_frame_offset: 10.0
463
+ len_traj_pred: ${len_traj_pred}
464
+ planner_mu_init:
465
+ - -0.1
466
+ - 0
467
+ - 0
468
+ - 0
469
+ - 0
470
+ planner_sigma_init:
471
+ - 0.02
472
+ - 0.1
473
+ - 0
474
+ - 0
475
+ - 0.3142
476
+ pose_fields:
477
+ - x
478
+ - 'y'
479
+ - z
480
+ - pitch
481
+ - yaw
482
+ pose_dim: 5
483
+ pos_dim: 3
484
+ angle_dim: 2
485
+ pose_mode: 3d
486
+ fps: 10.0
487
+ shared:
488
+ manifest_path: results/manifests/ai2thor/ai2thor_latent.json
489
+ backend:
490
+ _target_: far.data.backends.THORBackend
491
+ key_suffix: null
492
+ transform: null
493
+ sampler:
494
+ random_clip:
495
+ _target_: far.data.samplers.RandomClipSampler
496
+ clip_len: 16
497
+ frame_stride: 5
498
+ context_goal:
499
+ _target_: far.data.samplers.ContextGoalSampler
500
+ context_pool_size: 200
501
+ len_traj_pred: ${len_traj_pred}
502
+ goals_per_obs: ${goals_per_obs}
503
+ context_goal_test:
504
+ _target_: far.data.samplers.ContextGoalSampler
505
+ context_pool_size: 200
506
+ len_traj_pred: ${len_traj_pred}
507
+ goals_per_obs: 4
508
+ context_traj:
509
+ _target_: far.data.samplers.ContextTrajectorySampler
510
+ context_pool_size: 200
511
+ len_traj_pred: ${len_traj_pred}
512
+ formatter:
513
+ frame:
514
+ _target_: far.data.formatters.FrameFormatter
515
+ context_pose:
516
+ _target_: far.data.formatters.ContextPoseFormatter
517
+ spec: ${ai2thor_latent.spec}
518
+ train:
519
+ _target_: far.data.datasets.FormattedVideoDataset
520
+ formatter: ${ai2thor_latent.formatter.context_pose}
521
+ base_dataset:
522
+ _target_: far.data.datasets.VideoDataset
523
+ manifest_path: ${ai2thor_latent.shared.manifest_path}
524
+ index_path: results/indices/ai2thor/ai2thor_latent_train.json
525
+ backend: ${ai2thor_latent.shared.backend}
526
+ sampler: ${ai2thor_latent.sampler.context_goal}
527
+ transform: ${ai2thor_latent.shared.transform}
528
+ test:
529
+ _target_: far.data.datasets.FormattedVideoDataset
530
+ formatter: ${ai2thor_latent.formatter.context_pose}
531
+ base_dataset:
532
+ _target_: far.data.datasets.VideoDataset
533
+ manifest_path: ${ai2thor_latent.shared.manifest_path}
534
+ index_path: results/indices/ai2thor/ai2thor_latent_test.json
535
+ backend: ${ai2thor_latent.shared.backend}
536
+ sampler: ${ai2thor_latent.sampler.context_goal_test}
537
+ transform: ${ai2thor_latent.shared.transform}
538
+ model:
539
+ memory:
540
+ _target_: far.memories.pnp_cl.PnPRetrieval
541
+ ckpt_path: null
542
+ chunk_size: 10
543
+ num_sets: 4
544
+ softmax_temperature: 1.0
545
+ scoring: visual
546
+ metadata_hidden: 128
547
+ local_pos_stats: null
548
+ stride_to_seconds: 1.0
549
+ encoder_mode: frozen
550
+ adapter_hidden: 384
551
+ query_source: frame
552
+ query_conditioning: action
553
+ gate_eps: 0.05
554
+ query_vit:
555
+ _target_: far.memories.pnp_cl.QueryViT
556
+ img_height: 32
557
+ img_width: 32
558
+ patch_size: 2
559
+ in_chans: 4
560
+ embed_dim: 384
561
+ depth: 6
562
+ num_heads: 12
563
+ action_dim: ${action_dim}
564
+ key_dim: 256
565
+ mlp_ratio: 4.0
566
+ interact_dim: 4
567
+ wandb:
568
+ project: PMem-THOR
569
+ name: null
570
+ entity: null
571
+ mode: online
572
+ group: ai2thor_latent
573
+ tags:
574
+ - ai2thor_latent
575
+ - retriever_pretrain
576
+ - visual
577
+ run_name: retriever_thor_jepa_pretrain_ltp64
578
+ context_size: 4
579
+ action_dim: 5
580
+ pose_dim: 5
581
+ image_height: 256
582
+ image_width: 256
583
+ len_traj_pred: 64
584
+ goals_per_obs: 1
585
+ lr: 0.0002
586
+ weight_decay: 0.0
587
+ grad_clip_val: 1.0
588
+ iterations: 400000
589
+ warmup_steps: 5000
590
+ from_checkpoint: null
591
+ global_seed: 0
592
+ log_every: 100
593
+ ckpt_every: 25000
594
+ eval_every: 5000
595
+ bfloat16: false
596
+ batch_size: 16
597
+ num_workers: 12
598
+ datasets:
599
+ - ai2thor_latent
600
+ contrastive:
601
+ chunk_size: ${model.memory.chunk_size}
602
+ positive_pair_type:
603
+ - jepa
604
+ num_positives: 8
605
+ num_negatives: 24
606
+ cl_temperature: 0.1
607
+ cl_pose_radius: 0.5
608
+ cl_pose_fov_half_h: 180.0
609
+ cl_pose_fov_half_v: 180.0
610
+ cl_pose_num_samples: 1024
611
+ local_pos_stats: ${ai2thor_latent.spec.local_pos_stats}
bundles.json CHANGED
@@ -33,6 +33,20 @@
33
  "core_buffers"
34
  ]
35
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  {
37
  "corpus": "ai2thor_dyn",
38
  "arm": "temporal",
@@ -81,6 +95,20 @@
81
  "core_buffers"
82
  ]
83
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  {
85
  "corpus": "ai2thor_v3",
86
  "arm": "temporal",
@@ -177,6 +205,34 @@
177
  "core_buffers"
178
  ]
179
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
180
  {
181
  "corpus": "loopnav",
182
  "arm": "temporal",
@@ -241,6 +297,20 @@
241
  "core_buffers"
242
  ]
243
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
244
  {
245
  "corpus": "soundspaces_v1",
246
  "arm": "temporal",
 
33
  "core_buffers"
34
  ]
35
  },
36
+ {
37
+ "corpus": "ai2thor_dyn",
38
+ "arm": "retriever_object",
39
+ "step": "0012500",
40
+ "role": "retriever",
41
+ "release": true,
42
+ "path": "ai2thor_dyn/retriever_object/checkpoints/0012500.pth.tar",
43
+ "size_bytes": 66429235,
44
+ "sha256": "6d3d3e13be18d6e5d75bfb6ae59ef5e4965b11f48ae22133db917a99a44222f5",
45
+ "keys": [
46
+ "memory",
47
+ "pretrain_steps"
48
+ ]
49
+ },
50
  {
51
  "corpus": "ai2thor_dyn",
52
  "arm": "temporal",
 
95
  "core_buffers"
96
  ]
97
  },
98
+ {
99
+ "corpus": "ai2thor_v3",
100
+ "arm": "retriever_jepa",
101
+ "step": "0400000",
102
+ "role": "retriever",
103
+ "release": true,
104
+ "path": "ai2thor_v3/retriever_jepa/checkpoints/0400000.pth.tar",
105
+ "size_bytes": 66429235,
106
+ "sha256": "32ecb108654a7df2b3404c9508f4ada791c8c802b48558ab978eae5d9830119d",
107
+ "keys": [
108
+ "memory",
109
+ "pretrain_steps"
110
+ ]
111
+ },
112
  {
113
  "corpus": "ai2thor_v3",
114
  "arm": "temporal",
 
205
  "core_buffers"
206
  ]
207
  },
208
+ {
209
+ "corpus": "loopnav",
210
+ "arm": "retriever_jepa",
211
+ "step": "0400000",
212
+ "role": "retriever",
213
+ "release": true,
214
+ "path": "loopnav/retriever_jepa/checkpoints/0400000.pth.tar",
215
+ "size_bytes": 65900649,
216
+ "sha256": "41cf8aa4b4a274e9849d211b16a066b4d973f8763d26ed525483c90fe4306723",
217
+ "keys": [
218
+ "memory",
219
+ "pretrain_steps"
220
+ ]
221
+ },
222
+ {
223
+ "corpus": "loopnav",
224
+ "arm": "retriever_longlive",
225
+ "step": "0400000",
226
+ "role": "retriever",
227
+ "release": true,
228
+ "path": "loopnav/retriever_longlive/checkpoints/0400000.pth.tar",
229
+ "size_bytes": 65893225,
230
+ "sha256": "cb511f4b56853c7e06495151809a34d023c2db9f181f9d48c5d3e9f7d1a57af4",
231
+ "keys": [
232
+ "memory",
233
+ "pretrain_steps"
234
+ ]
235
+ },
236
  {
237
  "corpus": "loopnav",
238
  "arm": "temporal",
 
297
  "core_buffers"
298
  ]
299
  },
300
+ {
301
+ "corpus": "soundspaces_v1",
302
+ "arm": "retriever_audio",
303
+ "step": "0400000",
304
+ "role": "retriever",
305
+ "release": true,
306
+ "path": "soundspaces_v1/retriever_audio/checkpoints/0400000.pth.tar",
307
+ "size_bytes": 65900649,
308
+ "sha256": "0e792c5ee5d6f852f15e2d731d981b86e35bc055c65c92e9800c27ee063bf727",
309
+ "keys": [
310
+ "memory",
311
+ "pretrain_steps"
312
+ ]
313
+ },
314
  {
315
  "corpus": "soundspaces_v1",
316
  "arm": "temporal",
loopnav/retriever_jepa/checkpoints/0400000.pth.tar ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:41cf8aa4b4a274e9849d211b16a066b4d973f8763d26ed525483c90fe4306723
3
+ size 65900649
loopnav/retriever_jepa/config.yaml ADDED
@@ -0,0 +1,334 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ recon:
2
+ spec:
3
+ _target_: far.builders.data.build_dataspec
4
+ local_pos_stats:
5
+ min:
6
+ - -2.5
7
+ - -4
8
+ max:
9
+ - 5
10
+ - 4
11
+ meters_per_waypoint: 0.25
12
+ max_frame_offset: 128
13
+ len_traj_pred: ${len_traj_pred}
14
+ planner_mu_init:
15
+ - -0.1
16
+ - 0
17
+ - 0
18
+ planner_sigma_init:
19
+ - 0.02
20
+ - 0.1
21
+ - 0.3142
22
+ pose_fields:
23
+ - x
24
+ - 'y'
25
+ - z
26
+ - pitch
27
+ - yaw
28
+ pose_dim: 5
29
+ pos_dim: 2
30
+ angle_dim: 1
31
+ pose_mode: 2d
32
+ fps: 4.0
33
+ shared:
34
+ manifest_path: results/manifests/recon/recon_split.json
35
+ backend:
36
+ _target_: far.data.backends.NavigationBackend
37
+ transform:
38
+ _target_: far.data.transforms.make_centercrop_resize
39
+ image_height: ${image_height}
40
+ image_width: ${image_width}
41
+ sampler:
42
+ context_goal:
43
+ _target_: far.data.samplers.ContextGoalSampler
44
+ context_pool_size: 12
45
+ len_traj_pred: ${len_traj_pred}
46
+ goals_per_obs: ${goals_per_obs}
47
+ context_traj:
48
+ _target_: far.data.samplers.ContextTrajectorySampler
49
+ context_pool_size: 12
50
+ len_traj_pred: ${len_traj_pred}
51
+ formatter:
52
+ context_pose:
53
+ _target_: far.data.formatters.ContextPoseFormatter2D
54
+ spec: ${recon.spec}
55
+ train:
56
+ _target_: far.data.datasets.FormattedVideoDataset
57
+ formatter: ${recon.formatter.context_pose}
58
+ base_dataset:
59
+ _target_: far.data.datasets.VideoDataset
60
+ manifest_path: ${recon.shared.manifest_path}
61
+ index_path: results/indices/recon/recon_train.json
62
+ backend: ${recon.shared.backend}
63
+ sampler: ${recon.sampler.context_goal}
64
+ transform: ${recon.shared.transform}
65
+ test:
66
+ _target_: far.data.datasets.FormattedVideoDataset
67
+ formatter: ${recon.formatter.context_pose}
68
+ base_dataset:
69
+ _target_: far.data.datasets.VideoDataset
70
+ manifest_path: ${recon.shared.manifest_path}
71
+ index_path: results/indices/recon/recon_test.json
72
+ backend: ${recon.shared.backend}
73
+ sampler: ${recon.sampler.context_goal}
74
+ transform: ${recon.shared.transform}
75
+ eval:
76
+ _target_: far.data.datasets.FormattedVideoDataset
77
+ formatter: ${recon.formatter.context_pose}
78
+ base_dataset:
79
+ _target_: far.data.datasets.VideoDataset
80
+ manifest_path: ${recon.shared.manifest_path}
81
+ index_path: results/indices/recon/recon_test_time.json
82
+ backend: ${recon.shared.backend}
83
+ sampler: ${recon.sampler.context_traj}
84
+ transform: ${recon.shared.transform}
85
+ eval_traj:
86
+ _target_: far.data.datasets.FormattedVideoDataset
87
+ formatter: ${recon.formatter.context_pose}
88
+ base_dataset:
89
+ _target_: far.data.datasets.VideoDataset
90
+ manifest_path: ${recon.shared.manifest_path}
91
+ index_path: results/indices/recon/recon_test_plan.json
92
+ backend: ${recon.shared.backend}
93
+ sampler: ${recon.sampler.context_traj}
94
+ transform: ${recon.shared.transform}
95
+ loopnav:
96
+ spec:
97
+ _target_: far.builders.data.build_dataspec
98
+ local_pos_stats:
99
+ min:
100
+ - -2.5
101
+ - -4
102
+ - 0
103
+ max:
104
+ - 5
105
+ - 4
106
+ - 1
107
+ meters_per_waypoint: 0.25
108
+ max_frame_offset: 128
109
+ len_traj_pred: ${len_traj_pred}
110
+ planner_mu_init:
111
+ - -0.1
112
+ - 0
113
+ - 0
114
+ - 0
115
+ - 0
116
+ planner_sigma_init:
117
+ - 0.02
118
+ - 0.1
119
+ - 0
120
+ - 0
121
+ - 0.3142
122
+ pose_fields:
123
+ - x
124
+ - 'y'
125
+ - z
126
+ - pitch
127
+ - yaw
128
+ pose_dim: 5
129
+ pos_dim: 3
130
+ angle_dim: 2
131
+ pose_mode: 3d
132
+ fps: 20.0
133
+ shared:
134
+ manifest_path: results/manifests/loopnav/loopnav_split.json
135
+ backend:
136
+ _target_: far.data.backends.LoopNavBackend
137
+ transform:
138
+ _target_: far.data.transforms.make_resize
139
+ image_height: ${image_height}
140
+ image_width: ${image_width}
141
+ sampler:
142
+ random_clip:
143
+ _target_: far.data.samplers.RandomClipSampler
144
+ clip_len: 16
145
+ frame_stride: 10
146
+ context_goal:
147
+ _target_: far.data.samplers.ContextGoalSampler
148
+ context_pool_size: 100
149
+ len_traj_pred: ${len_traj_pred}
150
+ goals_per_obs: ${goals_per_obs}
151
+ context_traj:
152
+ _target_: far.data.samplers.ContextTrajectorySampler
153
+ context_pool_size: 100
154
+ len_traj_pred: ${len_traj_pred}
155
+ formatter:
156
+ frame:
157
+ _target_: far.data.formatters.FrameFormatter
158
+ context_pose:
159
+ _target_: far.data.formatters.ContextPoseFormatter
160
+ spec: ${loopnav.spec}
161
+ train_tokenizer:
162
+ _target_: far.data.datasets.FormattedIterableDataset
163
+ formatter: ${loopnav.formatter.frame}
164
+ base_dataset:
165
+ _target_: far.data.datasets.WebVideoIterableDataset
166
+ manifest_path: results/manifests/loopnav/loopnav.json
167
+ index_path: results/indices/loopnav/loopnav_train_ABCA.json
168
+ backend:
169
+ _target_: far.data.backends.WebDatasetBackend
170
+ sampler: ${loopnav.sampler.random_clip}
171
+ transform: ${loopnav.shared.transform}
172
+ cache_dir: null
173
+ cache_size: 100
174
+ shuffle_shards: 32
175
+ shuffle_samples: 100
176
+ loopnav_latent:
177
+ spec:
178
+ _target_: far.builders.data.build_dataspec
179
+ local_pos_stats:
180
+ min:
181
+ - 0
182
+ - 0
183
+ - 0
184
+ max:
185
+ - 1
186
+ - 1
187
+ - 1
188
+ meters_per_waypoint: 1.0
189
+ max_frame_offset: 20.0
190
+ len_traj_pred: ${len_traj_pred}
191
+ planner_mu_init:
192
+ - -0.1
193
+ - 0
194
+ - 0
195
+ - 0
196
+ - 0
197
+ planner_sigma_init:
198
+ - 0.02
199
+ - 0.1
200
+ - 0
201
+ - 0
202
+ - 0.3142
203
+ pose_fields:
204
+ - x
205
+ - 'y'
206
+ - z
207
+ - pitch
208
+ - yaw
209
+ pose_dim: 5
210
+ pos_dim: 3
211
+ angle_dim: 2
212
+ pose_mode: 3d
213
+ fps: 20.0
214
+ shared:
215
+ manifest_path: results/manifests/loopnav/loopnav_latent.json
216
+ backend:
217
+ _target_: far.data.backends.LatentLoopNavBackend
218
+ key_suffix: .keys.npy
219
+ transform: null
220
+ sampler:
221
+ random_clip:
222
+ _target_: far.data.samplers.RandomClipSampler
223
+ clip_len: 16
224
+ frame_stride: 10
225
+ context_goal:
226
+ _target_: far.data.samplers.ContextGoalSampler
227
+ context_pool_size: 200
228
+ len_traj_pred: ${len_traj_pred}
229
+ goals_per_obs: ${goals_per_obs}
230
+ goal_chunk_size: null
231
+ context_goal_test:
232
+ _target_: far.data.samplers.ContextGoalSampler
233
+ context_pool_size: 200
234
+ len_traj_pred: ${len_traj_pred}
235
+ goals_per_obs: 4
236
+ context_traj:
237
+ _target_: far.data.samplers.ContextTrajectorySampler
238
+ context_pool_size: 200
239
+ len_traj_pred: ${len_traj_pred}
240
+ formatter:
241
+ frame:
242
+ _target_: far.data.formatters.FrameFormatter
243
+ context_pose:
244
+ _target_: far.data.formatters.ContextPoseFormatter
245
+ spec: ${loopnav_latent.spec}
246
+ train:
247
+ _target_: far.data.datasets.FormattedVideoDataset
248
+ formatter: ${loopnav_latent.formatter.context_pose}
249
+ base_dataset:
250
+ _target_: far.data.datasets.VideoDataset
251
+ manifest_path: ${loopnav_latent.shared.manifest_path}
252
+ index_path: results/indices/loopnav/loopnav_latent_train.json
253
+ backend: ${loopnav_latent.shared.backend}
254
+ sampler: ${loopnav_latent.sampler.context_goal}
255
+ transform: ${loopnav_latent.shared.transform}
256
+ test:
257
+ _target_: far.data.datasets.FormattedVideoDataset
258
+ formatter: ${loopnav_latent.formatter.context_pose}
259
+ base_dataset:
260
+ _target_: far.data.datasets.VideoDataset
261
+ manifest_path: ${loopnav_latent.shared.manifest_path}
262
+ index_path: results/indices/loopnav/loopnav_latent_test.json
263
+ backend: ${loopnav_latent.shared.backend}
264
+ sampler: ${loopnav_latent.sampler.context_goal_test}
265
+ transform: ${loopnav_latent.shared.transform}
266
+ model:
267
+ memory:
268
+ _target_: far.memories.pnp_cl.PnPRetrieval
269
+ ckpt_path: null
270
+ chunk_size: 7
271
+ num_sets: 4
272
+ beam_size: 64
273
+ lambda_div: 0.0
274
+ softmax_temperature: 1.0
275
+ random_replace_prob: 0.0
276
+ random_replace_warmup_prob: 0.0
277
+ random_replace_warmup_steps: 50000
278
+ query_vit:
279
+ _target_: far.memories.pnp_cl.QueryViT
280
+ img_height: 18
281
+ img_width: 32
282
+ patch_size: 2
283
+ in_chans: 16
284
+ embed_dim: 384
285
+ depth: 6
286
+ num_heads: 12
287
+ action_dim: ${action_dim}
288
+ key_dim: 256
289
+ mlp_ratio: 4.0
290
+ query_adapter_hidden: null
291
+ wandb:
292
+ project: PMem-World
293
+ name: null
294
+ entity: null
295
+ mode: online
296
+ group: loopnav_latent
297
+ tags:
298
+ - loopnav_latent
299
+ - retriever_pretrain
300
+ run_name: retriever_jepa_pretrain
301
+ context_size: 4
302
+ action_dim: 5
303
+ pose_dim: 5
304
+ image_height: 360
305
+ image_width: 640
306
+ len_traj_pred: 128
307
+ goals_per_obs: 1
308
+ lr: 0.0002
309
+ weight_decay: 0.0
310
+ grad_clip_val: 1.0
311
+ iterations: 400000
312
+ warmup_steps: 5000
313
+ from_checkpoint: null
314
+ global_seed: 0
315
+ log_every: 100
316
+ ckpt_every: 50000
317
+ eval_every: 5000
318
+ bfloat16: false
319
+ batch_size: 64
320
+ num_workers: 12
321
+ datasets:
322
+ - loopnav_latent
323
+ contrastive:
324
+ chunk_size: ${model.memory.chunk_size}
325
+ positive_pair_type:
326
+ - jepa
327
+ num_positives: 8
328
+ num_negatives: 24
329
+ cl_temperature: 0.1
330
+ cl_pose_radius: 30.0
331
+ cl_pose_fov_half_h: 52.5
332
+ cl_pose_fov_half_v: 37.5
333
+ cl_pose_num_samples: 1024
334
+ local_pos_stats: ${loopnav_latent.spec.local_pos_stats}
loopnav/retriever_longlive/checkpoints/0400000.pth.tar ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cb511f4b56853c7e06495151809a34d023c2db9f181f9d48c5d3e9f7d1a57af4
3
+ size 65893225
loopnav/retriever_longlive/config.yaml ADDED
@@ -0,0 +1,440 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ recon:
2
+ spec:
3
+ _target_: far.builders.data.build_dataspec
4
+ local_pos_stats:
5
+ min:
6
+ - -2.5
7
+ - -4
8
+ max:
9
+ - 5
10
+ - 4
11
+ meters_per_waypoint: 0.25
12
+ max_frame_offset: 128
13
+ len_traj_pred: ${len_traj_pred}
14
+ planner_mu_init:
15
+ - -0.1
16
+ - 0
17
+ - 0
18
+ planner_sigma_init:
19
+ - 0.02
20
+ - 0.1
21
+ - 0.3142
22
+ pose_fields:
23
+ - x
24
+ - 'y'
25
+ - z
26
+ - pitch
27
+ - yaw
28
+ pose_dim: 5
29
+ pos_dim: 2
30
+ angle_dim: 1
31
+ pose_mode: 2d
32
+ fps: 4.0
33
+ shared:
34
+ manifest_path: results/manifests/recon/recon_split.json
35
+ backend:
36
+ _target_: far.data.backends.NavigationBackend
37
+ transform:
38
+ _target_: far.data.transforms.make_centercrop_resize
39
+ image_height: ${image_height}
40
+ image_width: ${image_width}
41
+ sampler:
42
+ context_goal:
43
+ _target_: far.data.samplers.ContextGoalSampler
44
+ context_pool_size: 12
45
+ len_traj_pred: ${len_traj_pred}
46
+ goals_per_obs: ${goals_per_obs}
47
+ context_traj:
48
+ _target_: far.data.samplers.ContextTrajectorySampler
49
+ context_pool_size: 12
50
+ len_traj_pred: ${len_traj_pred}
51
+ formatter:
52
+ context_pose:
53
+ _target_: far.data.formatters.ContextPoseFormatter2D
54
+ spec: ${recon.spec}
55
+ train:
56
+ _target_: far.data.datasets.FormattedVideoDataset
57
+ formatter: ${recon.formatter.context_pose}
58
+ base_dataset:
59
+ _target_: far.data.datasets.VideoDataset
60
+ manifest_path: ${recon.shared.manifest_path}
61
+ index_path: results/indices/recon/recon_train.json
62
+ backend: ${recon.shared.backend}
63
+ sampler: ${recon.sampler.context_goal}
64
+ transform: ${recon.shared.transform}
65
+ test:
66
+ _target_: far.data.datasets.FormattedVideoDataset
67
+ formatter: ${recon.formatter.context_pose}
68
+ base_dataset:
69
+ _target_: far.data.datasets.VideoDataset
70
+ manifest_path: ${recon.shared.manifest_path}
71
+ index_path: results/indices/recon/recon_test.json
72
+ backend: ${recon.shared.backend}
73
+ sampler: ${recon.sampler.context_goal}
74
+ transform: ${recon.shared.transform}
75
+ eval:
76
+ _target_: far.data.datasets.FormattedVideoDataset
77
+ formatter: ${recon.formatter.context_pose}
78
+ base_dataset:
79
+ _target_: far.data.datasets.VideoDataset
80
+ manifest_path: ${recon.shared.manifest_path}
81
+ index_path: results/indices/recon/recon_test_time.json
82
+ backend: ${recon.shared.backend}
83
+ sampler: ${recon.sampler.context_traj}
84
+ transform: ${recon.shared.transform}
85
+ eval_traj:
86
+ _target_: far.data.datasets.FormattedVideoDataset
87
+ formatter: ${recon.formatter.context_pose}
88
+ base_dataset:
89
+ _target_: far.data.datasets.VideoDataset
90
+ manifest_path: ${recon.shared.manifest_path}
91
+ index_path: results/indices/recon/recon_test_plan.json
92
+ backend: ${recon.shared.backend}
93
+ sampler: ${recon.sampler.context_traj}
94
+ transform: ${recon.shared.transform}
95
+ loopnav:
96
+ spec:
97
+ _target_: far.builders.data.build_dataspec
98
+ local_pos_stats:
99
+ min:
100
+ - -2.5
101
+ - -4
102
+ - 0
103
+ max:
104
+ - 5
105
+ - 4
106
+ - 1
107
+ meters_per_waypoint: 0.25
108
+ max_frame_offset: 128
109
+ len_traj_pred: ${len_traj_pred}
110
+ planner_mu_init:
111
+ - -0.1
112
+ - 0
113
+ - 0
114
+ - 0
115
+ - 0
116
+ planner_sigma_init:
117
+ - 0.02
118
+ - 0.1
119
+ - 0
120
+ - 0
121
+ - 0.3142
122
+ pose_fields:
123
+ - x
124
+ - 'y'
125
+ - z
126
+ - pitch
127
+ - yaw
128
+ pose_dim: 5
129
+ pos_dim: 3
130
+ angle_dim: 2
131
+ pose_mode: 3d
132
+ fps: 20.0
133
+ shared:
134
+ manifest_path: results/manifests/loopnav/loopnav_split.json
135
+ backend:
136
+ _target_: far.data.backends.LoopNavBackend
137
+ standard_frame: false
138
+ transform:
139
+ _target_: far.data.transforms.make_resize
140
+ image_height: ${image_height}
141
+ image_width: ${image_width}
142
+ sampler:
143
+ random_clip:
144
+ _target_: far.data.samplers.RandomClipSampler
145
+ clip_len: 16
146
+ frame_stride: 10
147
+ context_goal:
148
+ _target_: far.data.samplers.ContextGoalSampler
149
+ context_pool_size: 100
150
+ len_traj_pred: ${len_traj_pred}
151
+ goals_per_obs: ${goals_per_obs}
152
+ context_traj:
153
+ _target_: far.data.samplers.ContextTrajectorySampler
154
+ context_pool_size: 100
155
+ len_traj_pred: ${len_traj_pred}
156
+ formatter:
157
+ frame:
158
+ _target_: far.data.formatters.FrameFormatter
159
+ context_pose:
160
+ _target_: far.data.formatters.ContextPoseFormatter
161
+ spec: ${loopnav.spec}
162
+ train_tokenizer:
163
+ _target_: far.data.datasets.FormattedIterableDataset
164
+ formatter: ${loopnav.formatter.frame}
165
+ base_dataset:
166
+ _target_: far.data.datasets.WebVideoIterableDataset
167
+ manifest_path: results/manifests/loopnav/loopnav.json
168
+ index_path: results/indices/loopnav/loopnav_train_ABCA.json
169
+ backend:
170
+ _target_: far.data.backends.WebDatasetBackend
171
+ sampler: ${loopnav.sampler.random_clip}
172
+ transform: ${loopnav.shared.transform}
173
+ cache_dir: null
174
+ cache_size: 100
175
+ shuffle_shards: 32
176
+ shuffle_samples: 100
177
+ loopnav_latent:
178
+ spec:
179
+ _target_: far.builders.data.build_dataspec
180
+ local_pos_stats:
181
+ min:
182
+ - 0
183
+ - 0
184
+ - 0
185
+ max:
186
+ - 1
187
+ - 1
188
+ - 1
189
+ meters_per_waypoint: 1.0
190
+ max_frame_offset: 20.0
191
+ len_traj_pred: ${len_traj_pred}
192
+ planner_mu_init:
193
+ - -0.1
194
+ - 0
195
+ - 0
196
+ - 0
197
+ - 0
198
+ planner_sigma_init:
199
+ - 0.02
200
+ - 0.1
201
+ - 0
202
+ - 0
203
+ - 0.3142
204
+ pose_fields:
205
+ - x
206
+ - 'y'
207
+ - z
208
+ - pitch
209
+ - yaw
210
+ pose_dim: 5
211
+ pos_dim: 3
212
+ angle_dim: 2
213
+ pose_mode: 3d
214
+ fps: 20.0
215
+ shared:
216
+ manifest_path: results/manifests/loopnav/loopnav_latent.json
217
+ backend:
218
+ _target_: far.data.backends.LatentLoopNavBackend
219
+ key_suffix: null
220
+ standard_frame: false
221
+ transform: null
222
+ sampler:
223
+ random_clip:
224
+ _target_: far.data.samplers.RandomClipSampler
225
+ clip_len: 8
226
+ frame_stride: 10
227
+ context_goal:
228
+ _target_: far.data.samplers.ContextGoalSampler
229
+ context_pool_size: 200
230
+ len_traj_pred: ${len_traj_pred}
231
+ goals_per_obs: ${goals_per_obs}
232
+ context_goal_test:
233
+ _target_: far.data.samplers.ContextGoalSampler
234
+ context_pool_size: 200
235
+ len_traj_pred: ${len_traj_pred}
236
+ goals_per_obs: 4
237
+ context_traj:
238
+ _target_: far.data.samplers.ContextTrajectorySampler
239
+ context_pool_size: 200
240
+ len_traj_pred: ${len_traj_pred}
241
+ formatter:
242
+ frame:
243
+ _target_: far.data.formatters.FrameFormatter
244
+ context_pose:
245
+ _target_: far.data.formatters.ContextPoseFormatter
246
+ spec: ${loopnav_latent.spec}
247
+ train:
248
+ _target_: far.data.datasets.FormattedVideoDataset
249
+ formatter: ${loopnav_latent.formatter.context_pose}
250
+ base_dataset:
251
+ _target_: far.data.datasets.VideoDataset
252
+ manifest_path: ${loopnav_latent.shared.manifest_path}
253
+ index_path: results/indices/loopnav/loopnav_latent_train.json
254
+ backend: ${loopnav_latent.shared.backend}
255
+ sampler: ${loopnav_latent.sampler.context_goal}
256
+ transform: ${loopnav_latent.shared.transform}
257
+ test:
258
+ _target_: far.data.datasets.FormattedVideoDataset
259
+ formatter: ${loopnav_latent.formatter.context_pose}
260
+ base_dataset:
261
+ _target_: far.data.datasets.VideoDataset
262
+ manifest_path: ${loopnav_latent.shared.manifest_path}
263
+ index_path: results/indices/loopnav/loopnav_latent_test.json
264
+ backend: ${loopnav_latent.shared.backend}
265
+ sampler: ${loopnav_latent.sampler.context_goal_test}
266
+ transform: ${loopnav_latent.shared.transform}
267
+ soundspaces_latent:
268
+ spec:
269
+ _target_: far.builders.data.build_dataspec
270
+ local_pos_stats:
271
+ min:
272
+ - 0
273
+ - 0
274
+ - 0
275
+ max:
276
+ - 1
277
+ - 1
278
+ - 1
279
+ meters_per_waypoint: 1.0
280
+ max_frame_offset: 10.0
281
+ len_traj_pred: ${len_traj_pred}
282
+ planner_mu_init:
283
+ - -0.1
284
+ - 0
285
+ - 0
286
+ - 0
287
+ - 0
288
+ planner_sigma_init:
289
+ - 0.02
290
+ - 0.1
291
+ - 0
292
+ - 0
293
+ - 0.3142
294
+ pose_fields:
295
+ - x
296
+ - 'y'
297
+ - z
298
+ - pitch
299
+ - yaw
300
+ pose_dim: 5
301
+ pos_dim: 3
302
+ angle_dim: 2
303
+ pose_mode: 3d
304
+ fps: 10.0
305
+ shared:
306
+ manifest_path: results/manifests/soundspaces/soundspaces_latent.json
307
+ backend:
308
+ _target_: far.data.backends.SoundSpacesBackend
309
+ key_suffix: null
310
+ audio_mode: null
311
+ audio_window_hops: 48
312
+ transform: null
313
+ sampler:
314
+ random_clip:
315
+ _target_: far.data.samplers.RandomClipSampler
316
+ clip_len: 16
317
+ frame_stride: 5
318
+ context_goal:
319
+ _target_: far.data.samplers.ContextGoalSampler
320
+ context_pool_size: 200
321
+ len_traj_pred: ${len_traj_pred}
322
+ goals_per_obs: ${goals_per_obs}
323
+ context_goal_test:
324
+ _target_: far.data.samplers.ContextGoalSampler
325
+ context_pool_size: 200
326
+ len_traj_pred: ${len_traj_pred}
327
+ goals_per_obs: 4
328
+ context_traj:
329
+ _target_: far.data.samplers.ContextTrajectorySampler
330
+ context_pool_size: 200
331
+ len_traj_pred: ${len_traj_pred}
332
+ formatter:
333
+ frame:
334
+ _target_: far.data.formatters.FrameFormatter
335
+ context_pose:
336
+ _target_: far.data.formatters.ContextPoseFormatter
337
+ spec: ${soundspaces_latent.spec}
338
+ train:
339
+ _target_: far.data.datasets.FormattedVideoDataset
340
+ formatter: ${soundspaces_latent.formatter.context_pose}
341
+ base_dataset:
342
+ _target_: far.data.datasets.VideoDataset
343
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
344
+ index_path: results/indices/soundspaces/soundspaces_latent_train.json
345
+ backend: ${soundspaces_latent.shared.backend}
346
+ sampler: ${soundspaces_latent.sampler.context_goal}
347
+ transform: ${soundspaces_latent.shared.transform}
348
+ test:
349
+ _target_: far.data.datasets.FormattedVideoDataset
350
+ formatter: ${soundspaces_latent.formatter.context_pose}
351
+ base_dataset:
352
+ _target_: far.data.datasets.VideoDataset
353
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
354
+ index_path: results/indices/soundspaces/soundspaces_latent_test.json
355
+ backend: ${soundspaces_latent.shared.backend}
356
+ sampler: ${soundspaces_latent.sampler.context_goal_test}
357
+ transform: ${soundspaces_latent.shared.transform}
358
+ model:
359
+ memory:
360
+ _target_: far.memories.pnp_cl.PnPRetrieval
361
+ ckpt_path: null
362
+ chunk_size: 20
363
+ num_sets: 4
364
+ softmax_temperature: 1.0
365
+ scoring: visual
366
+ metadata_hidden: 128
367
+ local_pos_stats: null
368
+ stride_to_seconds: 1.0
369
+ encoder_mode: frozen
370
+ adapter_hidden: 384
371
+ query_source: frame
372
+ gate_eps: 0.05
373
+ query_vit:
374
+ _target_: far.memories.pnp_cl.QueryViT
375
+ img_height: 18
376
+ img_width: 32
377
+ patch_size: 2
378
+ in_chans: 16
379
+ embed_dim: 384
380
+ depth: 6
381
+ num_heads: 12
382
+ action_dim: ${action_dim}
383
+ key_dim: 256
384
+ mlp_ratio: 4.0
385
+ wandb:
386
+ project: PMem-World
387
+ name: null
388
+ entity: null
389
+ mode: online
390
+ group: loopnav_latent
391
+ tags:
392
+ - loopnav_latent
393
+ - retriever_pretrain
394
+ - longlive
395
+ run_name: retriever_longlive_pretrain
396
+ context_size: 4
397
+ action_dim: 5
398
+ pose_dim: 3
399
+ image_height: 224
400
+ image_width: 224
401
+ len_traj_pred: 128
402
+ goals_per_obs: 4
403
+ lr: 0.0003
404
+ weight_decay: 0.0001
405
+ grad_clip_val: 1.0
406
+ iterations: 400000
407
+ warmup_steps: 0
408
+ from_checkpoint: null
409
+ global_seed: 0
410
+ log_every: 100
411
+ ckpt_every: 25000
412
+ eval_every: 5000
413
+ bfloat16: false
414
+ batch_size: 32
415
+ num_workers: 12
416
+ datasets:
417
+ - loopnav_latent
418
+ contrastive:
419
+ chunk_size: ${model.memory.chunk_size}
420
+ positive_pair_type:
421
+ - temporal
422
+ - pose
423
+ - jepa
424
+ num_positives: 4
425
+ num_negatives: 16
426
+ cl_temperature: 0.1
427
+ cl_pose_radius: 30.0
428
+ cl_pose_fov_half_h: 52.5
429
+ cl_pose_fov_half_v: 37.5
430
+ cl_pose_num_samples: 1024
431
+ local_pos_stats: null
432
+ longlive:
433
+ hidden_dims:
434
+ - 64
435
+ - 128
436
+ - 256
437
+ delta_weight: 1.0
438
+ delta_margin: 0.85
439
+ delta_window: 3
440
+ smooth_weight: 1.0
soundspaces_v1/retriever_audio/checkpoints/0400000.pth.tar ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e792c5ee5d6f852f15e2d731d981b86e35bc055c65c92e9800c27ee063bf727
3
+ size 65900649
soundspaces_v1/retriever_audio/config.yaml ADDED
@@ -0,0 +1,429 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ recon:
2
+ spec:
3
+ _target_: far.builders.data.build_dataspec
4
+ local_pos_stats:
5
+ min:
6
+ - -2.5
7
+ - -4
8
+ max:
9
+ - 5
10
+ - 4
11
+ meters_per_waypoint: 0.25
12
+ max_frame_offset: 128
13
+ len_traj_pred: ${len_traj_pred}
14
+ planner_mu_init:
15
+ - -0.1
16
+ - 0
17
+ - 0
18
+ planner_sigma_init:
19
+ - 0.02
20
+ - 0.1
21
+ - 0.3142
22
+ pose_fields:
23
+ - x
24
+ - 'y'
25
+ - z
26
+ - pitch
27
+ - yaw
28
+ pose_dim: 5
29
+ pos_dim: 2
30
+ angle_dim: 1
31
+ pose_mode: 2d
32
+ fps: 4.0
33
+ shared:
34
+ manifest_path: results/manifests/recon/recon_split.json
35
+ backend:
36
+ _target_: far.data.backends.NavigationBackend
37
+ transform:
38
+ _target_: far.data.transforms.make_centercrop_resize
39
+ image_height: ${image_height}
40
+ image_width: ${image_width}
41
+ sampler:
42
+ context_goal:
43
+ _target_: far.data.samplers.ContextGoalSampler
44
+ context_pool_size: 12
45
+ len_traj_pred: ${len_traj_pred}
46
+ goals_per_obs: ${goals_per_obs}
47
+ context_traj:
48
+ _target_: far.data.samplers.ContextTrajectorySampler
49
+ context_pool_size: 12
50
+ len_traj_pred: ${len_traj_pred}
51
+ formatter:
52
+ context_pose:
53
+ _target_: far.data.formatters.ContextPoseFormatter2D
54
+ spec: ${recon.spec}
55
+ train:
56
+ _target_: far.data.datasets.FormattedVideoDataset
57
+ formatter: ${recon.formatter.context_pose}
58
+ base_dataset:
59
+ _target_: far.data.datasets.VideoDataset
60
+ manifest_path: ${recon.shared.manifest_path}
61
+ index_path: results/indices/recon/recon_train.json
62
+ backend: ${recon.shared.backend}
63
+ sampler: ${recon.sampler.context_goal}
64
+ transform: ${recon.shared.transform}
65
+ test:
66
+ _target_: far.data.datasets.FormattedVideoDataset
67
+ formatter: ${recon.formatter.context_pose}
68
+ base_dataset:
69
+ _target_: far.data.datasets.VideoDataset
70
+ manifest_path: ${recon.shared.manifest_path}
71
+ index_path: results/indices/recon/recon_test.json
72
+ backend: ${recon.shared.backend}
73
+ sampler: ${recon.sampler.context_goal}
74
+ transform: ${recon.shared.transform}
75
+ eval:
76
+ _target_: far.data.datasets.FormattedVideoDataset
77
+ formatter: ${recon.formatter.context_pose}
78
+ base_dataset:
79
+ _target_: far.data.datasets.VideoDataset
80
+ manifest_path: ${recon.shared.manifest_path}
81
+ index_path: results/indices/recon/recon_test_time.json
82
+ backend: ${recon.shared.backend}
83
+ sampler: ${recon.sampler.context_traj}
84
+ transform: ${recon.shared.transform}
85
+ eval_traj:
86
+ _target_: far.data.datasets.FormattedVideoDataset
87
+ formatter: ${recon.formatter.context_pose}
88
+ base_dataset:
89
+ _target_: far.data.datasets.VideoDataset
90
+ manifest_path: ${recon.shared.manifest_path}
91
+ index_path: results/indices/recon/recon_test_plan.json
92
+ backend: ${recon.shared.backend}
93
+ sampler: ${recon.sampler.context_traj}
94
+ transform: ${recon.shared.transform}
95
+ loopnav:
96
+ spec:
97
+ _target_: far.builders.data.build_dataspec
98
+ local_pos_stats:
99
+ min:
100
+ - -2.5
101
+ - -4
102
+ - 0
103
+ max:
104
+ - 5
105
+ - 4
106
+ - 1
107
+ meters_per_waypoint: 0.25
108
+ max_frame_offset: 128
109
+ len_traj_pred: ${len_traj_pred}
110
+ planner_mu_init:
111
+ - -0.1
112
+ - 0
113
+ - 0
114
+ - 0
115
+ - 0
116
+ planner_sigma_init:
117
+ - 0.02
118
+ - 0.1
119
+ - 0
120
+ - 0
121
+ - 0.3142
122
+ pose_fields:
123
+ - x
124
+ - 'y'
125
+ - z
126
+ - pitch
127
+ - yaw
128
+ pose_dim: 5
129
+ pos_dim: 3
130
+ angle_dim: 2
131
+ pose_mode: 3d
132
+ fps: 20.0
133
+ shared:
134
+ manifest_path: results/manifests/loopnav/loopnav_split.json
135
+ backend:
136
+ _target_: far.data.backends.LoopNavBackend
137
+ standard_frame: false
138
+ transform:
139
+ _target_: far.data.transforms.make_resize
140
+ image_height: ${image_height}
141
+ image_width: ${image_width}
142
+ sampler:
143
+ random_clip:
144
+ _target_: far.data.samplers.RandomClipSampler
145
+ clip_len: 16
146
+ frame_stride: 10
147
+ context_goal:
148
+ _target_: far.data.samplers.ContextGoalSampler
149
+ context_pool_size: 100
150
+ len_traj_pred: ${len_traj_pred}
151
+ goals_per_obs: ${goals_per_obs}
152
+ context_traj:
153
+ _target_: far.data.samplers.ContextTrajectorySampler
154
+ context_pool_size: 100
155
+ len_traj_pred: ${len_traj_pred}
156
+ formatter:
157
+ frame:
158
+ _target_: far.data.formatters.FrameFormatter
159
+ context_pose:
160
+ _target_: far.data.formatters.ContextPoseFormatter
161
+ spec: ${loopnav.spec}
162
+ train_tokenizer:
163
+ _target_: far.data.datasets.FormattedIterableDataset
164
+ formatter: ${loopnav.formatter.frame}
165
+ base_dataset:
166
+ _target_: far.data.datasets.WebVideoIterableDataset
167
+ manifest_path: results/manifests/loopnav/loopnav.json
168
+ index_path: results/indices/loopnav/loopnav_train_ABCA.json
169
+ backend:
170
+ _target_: far.data.backends.WebDatasetBackend
171
+ sampler: ${loopnav.sampler.random_clip}
172
+ transform: ${loopnav.shared.transform}
173
+ cache_dir: null
174
+ cache_size: 100
175
+ shuffle_shards: 32
176
+ shuffle_samples: 100
177
+ loopnav_latent:
178
+ spec:
179
+ _target_: far.builders.data.build_dataspec
180
+ local_pos_stats:
181
+ min:
182
+ - 0
183
+ - 0
184
+ - 0
185
+ max:
186
+ - 1
187
+ - 1
188
+ - 1
189
+ meters_per_waypoint: 1.0
190
+ max_frame_offset: 20.0
191
+ len_traj_pred: ${len_traj_pred}
192
+ planner_mu_init:
193
+ - -0.1
194
+ - 0
195
+ - 0
196
+ - 0
197
+ - 0
198
+ planner_sigma_init:
199
+ - 0.02
200
+ - 0.1
201
+ - 0
202
+ - 0
203
+ - 0.3142
204
+ pose_fields:
205
+ - x
206
+ - 'y'
207
+ - z
208
+ - pitch
209
+ - yaw
210
+ pose_dim: 5
211
+ pos_dim: 3
212
+ angle_dim: 2
213
+ pose_mode: 3d
214
+ fps: 20.0
215
+ shared:
216
+ manifest_path: results/manifests/loopnav/loopnav_latent.json
217
+ backend:
218
+ _target_: far.data.backends.LatentLoopNavBackend
219
+ key_suffix: .keys.npy
220
+ standard_frame: false
221
+ transform: null
222
+ sampler:
223
+ random_clip:
224
+ _target_: far.data.samplers.RandomClipSampler
225
+ clip_len: 16
226
+ frame_stride: 10
227
+ context_goal:
228
+ _target_: far.data.samplers.ContextGoalSampler
229
+ context_pool_size: 200
230
+ len_traj_pred: ${len_traj_pred}
231
+ goals_per_obs: ${goals_per_obs}
232
+ context_goal_test:
233
+ _target_: far.data.samplers.ContextGoalSampler
234
+ context_pool_size: 200
235
+ len_traj_pred: ${len_traj_pred}
236
+ goals_per_obs: 4
237
+ context_traj:
238
+ _target_: far.data.samplers.ContextTrajectorySampler
239
+ context_pool_size: 200
240
+ len_traj_pred: ${len_traj_pred}
241
+ formatter:
242
+ frame:
243
+ _target_: far.data.formatters.FrameFormatter
244
+ context_pose:
245
+ _target_: far.data.formatters.ContextPoseFormatter
246
+ spec: ${loopnav_latent.spec}
247
+ train:
248
+ _target_: far.data.datasets.FormattedVideoDataset
249
+ formatter: ${loopnav_latent.formatter.context_pose}
250
+ base_dataset:
251
+ _target_: far.data.datasets.VideoDataset
252
+ manifest_path: ${loopnav_latent.shared.manifest_path}
253
+ index_path: results/indices/loopnav/loopnav_latent_train.json
254
+ backend: ${loopnav_latent.shared.backend}
255
+ sampler: ${loopnav_latent.sampler.context_goal}
256
+ transform: ${loopnav_latent.shared.transform}
257
+ test:
258
+ _target_: far.data.datasets.FormattedVideoDataset
259
+ formatter: ${loopnav_latent.formatter.context_pose}
260
+ base_dataset:
261
+ _target_: far.data.datasets.VideoDataset
262
+ manifest_path: ${loopnav_latent.shared.manifest_path}
263
+ index_path: results/indices/loopnav/loopnav_latent_test.json
264
+ backend: ${loopnav_latent.shared.backend}
265
+ sampler: ${loopnav_latent.sampler.context_goal_test}
266
+ transform: ${loopnav_latent.shared.transform}
267
+ soundspaces_latent:
268
+ spec:
269
+ _target_: far.builders.data.build_dataspec
270
+ local_pos_stats:
271
+ min:
272
+ - 0
273
+ - 0
274
+ - 0
275
+ max:
276
+ - 1
277
+ - 1
278
+ - 1
279
+ meters_per_waypoint: 1.0
280
+ max_frame_offset: 10.0
281
+ len_traj_pred: ${len_traj_pred}
282
+ planner_mu_init:
283
+ - -0.1
284
+ - 0
285
+ - 0
286
+ - 0
287
+ - 0
288
+ planner_sigma_init:
289
+ - 0.02
290
+ - 0.1
291
+ - 0
292
+ - 0
293
+ - 0.3142
294
+ pose_fields:
295
+ - x
296
+ - 'y'
297
+ - z
298
+ - pitch
299
+ - yaw
300
+ pose_dim: 5
301
+ pos_dim: 3
302
+ angle_dim: 2
303
+ pose_mode: 3d
304
+ fps: 10.0
305
+ shared:
306
+ manifest_path: results/manifests/soundspaces/soundspaces_latent.json
307
+ backend:
308
+ _target_: far.data.backends.SoundSpacesBackend
309
+ key_suffix: null
310
+ audio_mode: frames
311
+ audio_window_hops: 48
312
+ transform: null
313
+ sampler:
314
+ random_clip:
315
+ _target_: far.data.samplers.RandomClipSampler
316
+ clip_len: 16
317
+ frame_stride: 5
318
+ context_goal:
319
+ _target_: far.data.samplers.ContextGoalSampler
320
+ context_pool_size: 200
321
+ len_traj_pred: ${len_traj_pred}
322
+ goals_per_obs: ${goals_per_obs}
323
+ context_goal_test:
324
+ _target_: far.data.samplers.ContextGoalSampler
325
+ context_pool_size: 200
326
+ len_traj_pred: ${len_traj_pred}
327
+ goals_per_obs: 4
328
+ context_traj:
329
+ _target_: far.data.samplers.ContextTrajectorySampler
330
+ context_pool_size: 200
331
+ len_traj_pred: ${len_traj_pred}
332
+ formatter:
333
+ frame:
334
+ _target_: far.data.formatters.FrameFormatter
335
+ context_pose:
336
+ _target_: far.data.formatters.ContextPoseFormatter
337
+ spec: ${soundspaces_latent.spec}
338
+ train:
339
+ _target_: far.data.datasets.FormattedVideoDataset
340
+ formatter: ${soundspaces_latent.formatter.context_pose}
341
+ base_dataset:
342
+ _target_: far.data.datasets.VideoDataset
343
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
344
+ index_path: results/indices/soundspaces/soundspaces_latent_train.json
345
+ backend: ${soundspaces_latent.shared.backend}
346
+ sampler: ${soundspaces_latent.sampler.context_goal}
347
+ transform: ${soundspaces_latent.shared.transform}
348
+ test:
349
+ _target_: far.data.datasets.FormattedVideoDataset
350
+ formatter: ${soundspaces_latent.formatter.context_pose}
351
+ base_dataset:
352
+ _target_: far.data.datasets.VideoDataset
353
+ manifest_path: ${soundspaces_latent.shared.manifest_path}
354
+ index_path: results/indices/soundspaces/soundspaces_latent_test.json
355
+ backend: ${soundspaces_latent.shared.backend}
356
+ sampler: ${soundspaces_latent.sampler.context_goal_test}
357
+ transform: ${soundspaces_latent.shared.transform}
358
+ model:
359
+ memory:
360
+ _target_: far.memories.pnp_cl.PnPRetrieval
361
+ ckpt_path: null
362
+ chunk_size: 10
363
+ num_sets: 4
364
+ softmax_temperature: 1.0
365
+ scoring: visual
366
+ metadata_hidden: 128
367
+ local_pos_stats: null
368
+ stride_to_seconds: 1.0
369
+ encoder_mode: frozen
370
+ adapter_hidden: 384
371
+ query_source: frame
372
+ gate_eps: 0.05
373
+ query_vit:
374
+ _target_: far.memories.pnp_cl.QueryViT
375
+ img_height: 64
376
+ img_width: 48
377
+ patch_size: 4
378
+ in_chans: 4
379
+ embed_dim: 384
380
+ depth: 6
381
+ num_heads: 12
382
+ action_dim: ${action_dim}
383
+ key_dim: 256
384
+ mlp_ratio: 4.0
385
+ wandb:
386
+ project: PMem-SS
387
+ name: null
388
+ entity: null
389
+ mode: online
390
+ group: soundspaces_latent
391
+ tags:
392
+ - soundspaces_latent
393
+ - retriever_pretrain
394
+ - audio
395
+ run_name: retriever_audio_jepa_pretrain_ltp64
396
+ context_size: 4
397
+ action_dim: 5
398
+ pose_dim: 5
399
+ image_height: 256
400
+ image_width: 256
401
+ len_traj_pred: 64
402
+ goals_per_obs: 1
403
+ lr: 0.0002
404
+ weight_decay: 0.0
405
+ grad_clip_val: 1.0
406
+ iterations: 400000
407
+ warmup_steps: 5000
408
+ from_checkpoint: null
409
+ global_seed: 0
410
+ log_every: 100
411
+ ckpt_every: 25000
412
+ eval_every: 5000
413
+ bfloat16: false
414
+ batch_size: 16
415
+ num_workers: 12
416
+ datasets:
417
+ - soundspaces_latent
418
+ contrastive:
419
+ chunk_size: ${model.memory.chunk_size}
420
+ positive_pair_type:
421
+ - jepa
422
+ num_positives: 8
423
+ num_negatives: 24
424
+ cl_temperature: 0.1
425
+ cl_pose_radius: 0.5
426
+ cl_pose_fov_half_h: 180.0
427
+ cl_pose_fov_half_v: 180.0
428
+ cl_pose_num_samples: 1024
429
+ local_pos_stats: ${soundspaces_latent.spec.local_pos_stats}