Spaces:
Running on CPU Upgrade
Running on CPU Upgrade
Deploy HF RL Explorer
Browse files- DESIGN.md +30 -0
- README.md +13 -5
- app/dockerfile.py +1 -1
- app/http.py +1 -0
- app/main.py +7 -2
- app/seo.py +197 -48
- app/seo_tasks.py +115 -0
- app/space_checks.py +13 -0
- web/index.html +1 -1
- web/js/seo-meta.js +17 -0
- web/js/space.js +17 -3
- web/js/util.js +26 -5
- web/rlx.css +1 -0
- web/social/rl-explorer.png +0 -0
DESIGN.md
CHANGED
|
@@ -242,3 +242,33 @@ This public tooling directory contains only the Explorer. The dashboard applicat
|
|
| 242 |
its assets, membership checks, tests and deployment tooling are maintained separately
|
| 243 |
in the private `FineEnvs/RL-Explorer-admin` Space. The Explorer reads shared settings
|
| 244 |
and moderation requests from the data bucket without importing the admin app.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 242 |
its assets, membership checks, tests and deployment tooling are maintained separately
|
| 243 |
in the private `FineEnvs/RL-Explorer-admin` Space. The Explorer reads shared settings
|
| 244 |
and moderation requests from the data bucket without importing the admin app.
|
| 245 |
+
|
| 246 |
+
## Search indexing and social previews
|
| 247 |
+
|
| 248 |
+
The Space card uses the committed `web/social/rl-explorer.png` thumbnail, so a
|
| 249 |
+
sharing crawler does not have to start the application. HTML pages also publish
|
| 250 |
+
Open Graph and Twitter metadata with 1200Γ630 previews. Page-specific images fit
|
| 251 |
+
long names into the card; public task links retain their own canonical address
|
| 252 |
+
both in the initial HTML and after browser navigation.
|
| 253 |
+
|
| 254 |
+
`/sitemap.xml` links to shards of up to 10,000 public task URLs. Dataset tasks come
|
| 255 |
+
from local indexes or the immutable catalog snapshot, with MiMo entries deduplicated.
|
| 256 |
+
OpenEnv task coordinates come from advertised split counts on checked public
|
| 257 |
+
Spaces. Their sitemaps expand ranges one shard at a time, without downloading tasks
|
| 258 |
+
or creating millions of URLs in memory. These counts describe published entries,
|
| 259 |
+
not individually tested episodes or Google-indexed pages.
|
| 260 |
+
|
| 261 |
+
Task HTML includes public task text or structured input facts. Public row datasets
|
| 262 |
+
and OpenEnv Task API records can be read anonymously on demand, with four concurrent
|
| 263 |
+
reads and bounded caches. These reads do not build an index, wake a Space, run an
|
| 264 |
+
episode, or forward a visitor's token. The existing answer-withholding rules apply.
|
| 265 |
+
Invalid tasks return 404; temporary upstream failures return 503 with Retry-After.
|
| 266 |
+
Private, gated and hidden content is excluded from sitemap generation.
|
| 267 |
+
|
| 268 |
+
Public read APIs may be fetched to render the interactive pages, but API responses
|
| 269 |
+
carry X-Robots-Tag: noindex. Rollouts, account pages and comparisons remain excluded.
|
| 270 |
+
Submit `https://fineenvs-rl-explorer.hf.space/sitemap.xml` for the matching URL-prefix
|
| 271 |
+
property in Google Search Console to monitor discovery and indexing. Robots.txt
|
| 272 |
+
also advertises it. Google chooses which pages to index; no application setting
|
| 273 |
+
can guarantee indexing or rankings. See Google's sitemap guidance:
|
| 274 |
+
https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap
|
README.md
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
---
|
| 2 |
-
title:
|
| 3 |
emoji: π€
|
| 4 |
colorFrom: yellow
|
| 5 |
colorTo: gray
|
|
@@ -7,7 +7,8 @@ sdk: docker
|
|
| 7 |
app_port: 7860
|
| 8 |
pinned: false
|
| 9 |
license: apache-2.0
|
| 10 |
-
short_description:
|
|
|
|
| 11 |
hf_oauth: true
|
| 12 |
hf_oauth_expiration_minutes: 1440
|
| 13 |
hf_oauth_scopes:
|
|
@@ -29,10 +30,17 @@ tags:
|
|
| 29 |
- mcp
|
| 30 |
---
|
| 31 |
|
| 32 |
-
#
|
| 33 |
|
| 34 |
-
|
| 35 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
- **Harbor datasets**: every task folder indexed (instruction, tests, image, grader, multi-step tasks, resources);
|
| 38 |
run one with OpenCode, Terminus 2, mini-SWE-agent or Pi on an HF Sandbox, graded by the task's own tests.
|
|
|
|
| 1 |
---
|
| 2 |
+
title: RL Environments on Hugging Face
|
| 3 |
emoji: π€
|
| 4 |
colorFrom: yellow
|
| 5 |
colorTo: gray
|
|
|
|
| 7 |
app_port: 7860
|
| 8 |
pinned: false
|
| 9 |
license: apache-2.0
|
| 10 |
+
short_description: Explore RL environments, browse tasks and run agent rollouts
|
| 11 |
+
thumbnail: https://huggingface.co/spaces/FineEnvs/RL-Explorer/resolve/main/web/social/rl-explorer.png
|
| 12 |
hf_oauth: true
|
| 13 |
hf_oauth_expiration_minutes: 1440
|
| 14 |
hf_oauth_scopes:
|
|
|
|
| 30 |
- mcp
|
| 31 |
---
|
| 32 |
|
| 33 |
+
# RL environments on the Hugging Face Hub
|
| 34 |
|
| 35 |
+
**HF RL Explorer** helps you discover reinforcement learning environments on the Hugging Face Hub.
|
| 36 |
+
Explore public datasets and checked environment Spaces across OpenEnv, Harbor, MiMo, NeMo Gym and Verifiers.
|
| 37 |
+
Inspect individual tasks, tools and reward functions, then run supported agent rollouts and compare results.
|
| 38 |
+
|
| 39 |
+
[Open the RL environment explorer](https://fineenvs-rl-explorer.hf.space/) Β·
|
| 40 |
+
[Explore FineEnvs environments](https://fineenvs-rl-explorer.hf.space/?owner=FineEnvs) Β·
|
| 41 |
+
[Public task sitemap](https://fineenvs-rl-explorer.hf.space/sitemap.xml)
|
| 42 |
+
|
| 43 |
+

|
| 44 |
|
| 45 |
- **Harbor datasets**: every task folder indexed (instruction, tests, image, grader, multi-step tasks, resources);
|
| 46 |
run one with OpenCode, Terminus 2, mini-SWE-agent or Pi on an HF Sandbox, graded by the task's own tests.
|
app/dockerfile.py
CHANGED
|
@@ -13,7 +13,7 @@ import re
|
|
| 13 |
import shlex
|
| 14 |
from dataclasses import dataclass, field
|
| 15 |
|
| 16 |
-
IGNORED = {"CMD", "ENTRYPOINT", "EXPOSE", "LABEL", "HEALTHCHECK", "VOLUME", "STOPSIGNAL", "MAINTAINER"
|
| 17 |
|
| 18 |
|
| 19 |
@dataclass
|
|
|
|
| 13 |
import shlex
|
| 14 |
from dataclasses import dataclass, field
|
| 15 |
|
| 16 |
+
IGNORED = {"CMD", "ENTRYPOINT", "EXPOSE", "LABEL", "HEALTHCHECK", "VOLUME", "STOPSIGNAL", "MAINTAINER"}
|
| 17 |
|
| 18 |
|
| 19 |
@dataclass
|
app/http.py
CHANGED
|
@@ -79,6 +79,7 @@ def install(app: FastAPI) -> None:
|
|
| 79 |
# APIs can contain a visitor's private dataset, run, or account. A shared
|
| 80 |
# proxy or the browser's cache must never reuse them for another visitor.
|
| 81 |
resp.headers["Cache-Control"] = "private, no-store"
|
|
|
|
| 82 |
else:
|
| 83 |
resp.headers.setdefault("Cache-Control", "no-cache")
|
| 84 |
return resp
|
|
|
|
| 79 |
# APIs can contain a visitor's private dataset, run, or account. A shared
|
| 80 |
# proxy or the browser's cache must never reuse them for another visitor.
|
| 81 |
resp.headers["Cache-Control"] = "private, no-store"
|
| 82 |
+
resp.headers["X-Robots-Tag"] = "noindex"
|
| 83 |
else:
|
| 84 |
resp.headers.setdefault("Cache-Control", "no-cache")
|
| 85 |
return resp
|
app/main.py
CHANGED
|
@@ -809,16 +809,21 @@ def sitemap_tasks(n: int, request: Request):
|
|
| 809 |
return seo.sitemap_tasks(request, n)
|
| 810 |
|
| 811 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 812 |
@app.get("/og.png", include_in_schema=False)
|
| 813 |
def og_site():
|
| 814 |
return seo.og_image("/")
|
| 815 |
|
| 816 |
|
| 817 |
@app.get("/og/{path:path}", include_in_schema=False)
|
| 818 |
-
def og_page(path: str):
|
| 819 |
if not path.endswith(".png") or len(path) > 600:
|
| 820 |
raise HTTPException(404, "no such image")
|
| 821 |
-
return seo.og_image("/" + path[:-4])
|
| 822 |
|
| 823 |
|
| 824 |
@app.exception_handler(404)
|
|
|
|
| 809 |
return seo.sitemap_tasks(request, n)
|
| 810 |
|
| 811 |
|
| 812 |
+
@app.get("/sitemap-space-tasks-{n}.xml", include_in_schema=False)
|
| 813 |
+
def sitemap_space_tasks(n: int, request: Request):
|
| 814 |
+
return seo.sitemap_space_tasks(request, n)
|
| 815 |
+
|
| 816 |
+
|
| 817 |
@app.get("/og.png", include_in_schema=False)
|
| 818 |
def og_site():
|
| 819 |
return seo.og_image("/")
|
| 820 |
|
| 821 |
|
| 822 |
@app.get("/og/{path:path}", include_in_schema=False)
|
| 823 |
+
def og_page(path: str, request: Request):
|
| 824 |
if not path.endswith(".png") or len(path) > 600:
|
| 825 |
raise HTTPException(404, "no such image")
|
| 826 |
+
return seo.og_image("/" + path[:-4], request.query_params)
|
| 827 |
|
| 828 |
|
| 829 |
@app.exception_handler(404)
|
app/seo.py
CHANGED
|
@@ -4,9 +4,9 @@ data (a Dataset for an environment, a SoftwareApplication for a Space, a task as
|
|
| 4 |
in its body, the page's content as plain HTML, for crawlers that don't run JavaScript (the app replaces it as it
|
| 5 |
starts). Plus robots.txt, a sitemap of every environment and every indexed task, and a preview image per page.
|
| 6 |
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
"""
|
| 11 |
|
| 12 |
from __future__ import annotations
|
|
@@ -19,22 +19,22 @@ import threading
|
|
| 19 |
import time
|
| 20 |
from functools import lru_cache
|
| 21 |
from typing import Any
|
| 22 |
-
from urllib.parse import quote
|
| 23 |
|
| 24 |
from fastapi import Request
|
| 25 |
from fastapi.responses import HTMLResponse, PlainTextResponse, Response
|
| 26 |
|
| 27 |
-
from . import catalog, config
|
| 28 |
|
| 29 |
SITE = "HF RL Explorer"
|
| 30 |
-
TAGLINE = "
|
| 31 |
-
DESCRIPTION = ("Explore reinforcement learning environments on Hugging Face
|
| 32 |
-
"
|
| 33 |
KIND = {"harbor": "Harbor dataset", "verifiers": "Verifiers environment", "nemo-gym": "NeMo Gym dataset", "rows": "RL dataset",
|
| 34 |
"mimo": "MiMo RL release", "openenv": "OpenEnv Space", "space": "environment Space"}
|
| 35 |
MIMO = "XiaomiMiMo/MiMo-V2.6-RL-oss"
|
| 36 |
NOINDEX = re.compile(r"^/(run/|runs$|compare/)")
|
| 37 |
-
SITEMAP_CHUNK =
|
| 38 |
|
| 39 |
|
| 40 |
# ββ where we are βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
|
@@ -92,6 +92,15 @@ def listing() -> tuple[dict[str, dict[str, Any]], list[dict[str, Any]]]:
|
|
| 92 |
return by, rows
|
| 93 |
|
| 94 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
def index_rows(spec: str) -> list[dict[str, Any]]:
|
| 96 |
"""An environment's tasks, if they are already known here (never built for a crawler): [{path, title, brief, category}]."""
|
| 97 |
if spec == MIMO:
|
|
@@ -100,7 +109,10 @@ def index_rows(spec: str) -> list[dict[str, Any]]:
|
|
| 100 |
p = catalog._index_path(spec)
|
| 101 |
mtime = p.stat().st_mtime
|
| 102 |
except (OSError, ValueError):
|
| 103 |
-
|
|
|
|
|
|
|
|
|
|
| 104 |
return _harbor_rows(spec, mtime)
|
| 105 |
|
| 106 |
|
|
@@ -114,9 +126,17 @@ def _mimo_rows() -> list[dict[str, Any]]:
|
|
| 114 |
@lru_cache(maxsize=32)
|
| 115 |
def _harbor_rows(spec: str, mtime: float) -> list[dict[str, Any]]:
|
| 116 |
idx = catalog._read_index(spec)
|
|
|
|
|
|
|
| 117 |
return [{"path": t["path"], "title": t.get("title"), "brief": t.get("brief"), "category": t.get("category")} for t in (idx or {}).get("tasks") or []]
|
| 118 |
|
| 119 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 120 |
def display_name(row: dict[str, Any] | None, spec: str) -> str:
|
| 121 |
"""What to call an environment: its card's heading when that names it, else the repository name with the heading."""
|
| 122 |
repo = spec.split("/")[-1]
|
|
@@ -145,9 +165,10 @@ def page(request: Request, *, title: str, description: str, path: str, body: str
|
|
| 145 |
base = base_url(request)
|
| 146 |
url = base + path
|
| 147 |
full = title if SITE in title else f"{title} Β· {SITE}"
|
| 148 |
-
img = base + (image or "/
|
| 149 |
head = "\n".join([
|
| 150 |
f'<link rel="canonical" href="{esc(url)}">',
|
|
|
|
| 151 |
'<meta name="robots" content="noindex, follow">' if noindex else '<meta name="robots" content="index, follow, max-image-preview:large, max-snippet:-1">',
|
| 152 |
f'<meta property="og:site_name" content="{SITE}">',
|
| 153 |
f'<meta property="og:type" content="{og_type}">',
|
|
@@ -155,6 +176,8 @@ def page(request: Request, *, title: str, description: str, path: str, body: str
|
|
| 155 |
f'<meta property="og:description" content="{esc(description)}">',
|
| 156 |
f'<meta property="og:url" content="{esc(url)}">',
|
| 157 |
f'<meta property="og:image" content="{esc(img)}">',
|
|
|
|
|
|
|
| 158 |
'<meta property="og:image:width" content="1200">',
|
| 159 |
'<meta property="og:image:height" content="630">',
|
| 160 |
f'<meta property="og:image:alt" content="{esc(full)}">',
|
|
@@ -162,6 +185,7 @@ def page(request: Request, *, title: str, description: str, path: str, body: str
|
|
| 162 |
f'<meta name="twitter:title" content="{esc(full)}">',
|
| 163 |
f'<meta name="twitter:description" content="{esc(description)}">',
|
| 164 |
f'<meta name="twitter:image" content="{esc(img)}">',
|
|
|
|
| 165 |
*[f'<script type="application/ld+json">{_ld(x)}</script>' for x in jsonld or []],
|
| 166 |
])
|
| 167 |
doc = template()
|
|
@@ -170,7 +194,7 @@ def page(request: Request, *, title: str, description: str, path: str, body: str
|
|
| 170 |
doc = doc.replace("</head>", f"{head}\n</head>", 1)
|
| 171 |
if body: # the page as plain HTML until the app starts (crawlers that don't run JavaScript read this)
|
| 172 |
doc = re.sub(r'<main id="view">(.*?)</main>', lambda m: f'<main id="view">{m.group(1)}<div class="wrap page ssr">{body}</div></main>', doc, count=1, flags=re.S)
|
| 173 |
-
return HTMLResponse(doc, status_code=status, headers={"Cache-Control": "no-cache"})
|
| 174 |
|
| 175 |
|
| 176 |
def _ld(x: dict[str, Any]) -> str:
|
|
@@ -191,18 +215,15 @@ def crumbs(base: str, items: list[tuple[str, str]]) -> tuple[str, dict[str, Any]
|
|
| 191 |
def home(request: Request) -> HTMLResponse:
|
| 192 |
base = base_url(request)
|
| 193 |
by, rows = listing()
|
|
|
|
| 194 |
top = sorted(rows, key=lambda r: (r.get("trending") or 0) * 1e9 + (r.get("downloads") or 0), reverse=True)[:120]
|
| 195 |
-
|
| 196 |
-
sp = sum(1 for r in rows if r["kind"] == "space")
|
| 197 |
-
tasks = sum((r.get("indexed") or {}).get("tasks") or 0 for r in rows)
|
| 198 |
-
desc = (f"{ds:,} RL environment datasets and {sp:,} environment Spaces on Hugging Face, across OpenEnv, Harbor, Verifiers, "
|
| 199 |
-
f"NeMo Gym and more: see what each task asks and how it's graded, then run an agent on it." if rows else DESCRIPTION)
|
| 200 |
body = (f"<header class=\"tp-head\"><h1>{TAGLINE}</h1><p class=\"lede\">{esc(desc)}</p></header>"
|
| 201 |
+ _env_list(top) + '<p><a href="/community">Community rollouts</a> Β· <a href="/d/XiaomiMiMo/MiMo-V2.6-RL-oss">MiMo-V2.6 RL</a></p>')
|
| 202 |
ld = [{"@type": "WebSite", "name": SITE, "alternateName": TAGLINE, "url": base + "/", "description": desc,
|
| 203 |
"potentialAction": {"@type": "SearchAction", "target": {"@type": "EntryPoint", "urlTemplate": base + "/?q={search_term_string}"},
|
| 204 |
"query-input": "required name=search_term_string"},
|
| 205 |
-
"publisher": {"@type": "Organization", "name": "
|
| 206 |
{"@type": "CollectionPage", "name": TAGLINE, "url": base + "/", "about": "Reinforcement learning environments",
|
| 207 |
"mainEntity": {"@type": "ItemList", "numberOfItems": len(top), "itemListElement": [
|
| 208 |
{"@type": "ListItem", "position": i + 1, "url": base + _href(r), "name": r.get("heading") or r["id"]} for i, r in enumerate(top[:50])]}}]
|
|
@@ -221,10 +242,10 @@ def _env_list(rows: list[dict[str, Any]]) -> str:
|
|
| 221 |
def environment(request: Request, spec: str) -> HTMLResponse:
|
| 222 |
base = base_url(request)
|
| 223 |
by, _ = listing()
|
| 224 |
-
r = by.
|
| 225 |
name = display_name(r, spec)
|
| 226 |
kind = kind_of(r, spec)
|
| 227 |
-
tasks = index_rows(spec) if r
|
| 228 |
n = len(tasks) or ((r or {}).get("indexed") or {}).get("tasks")
|
| 229 |
brief = clip((r or {}).get("brief") or "", 300)
|
| 230 |
desc = clip(f"{name}: {kind} on Hugging Face{f' with {n:,} tasks' if n else ''}. "
|
|
@@ -243,16 +264,24 @@ def environment(request: Request, spec: str) -> HTMLResponse:
|
|
| 243 |
"includedInDataCatalog": {"@type": "DataCatalog", "name": SITE, "url": base + "/"},
|
| 244 |
"distribution": [{"@type": "DataDownload", "encodingFormat": "application/octet-stream", "contentUrl": f"https://huggingface.co/datasets/{spec}"}]},
|
| 245 |
cld]
|
| 246 |
-
return page(request, title=f"{name} Β· {kind}", description=desc, path=path, body=body, jsonld=ld, image=f"/og{path}.png" if
|
| 247 |
|
| 248 |
|
| 249 |
def task(request: Request, spec: str, ref: str) -> HTMLResponse:
|
| 250 |
base = base_url(request)
|
| 251 |
by, _ = listing()
|
| 252 |
-
r = by.
|
| 253 |
env_name = display_name(r, spec)
|
| 254 |
-
known = r is not None
|
| 255 |
row = next((t for t in index_rows(spec) if t["path"] == ref), None) if known else None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 256 |
split, _, n = ref.rpartition("/")
|
| 257 |
# a row of a dataset read as rows (no index here): "Row 3 (train)" rather than a bare "3"
|
| 258 |
fallback = f"Row {n} ({split})" if split and n.isdigit() and "/" not in split else ref.rsplit("/", 1)[-1]
|
|
@@ -262,18 +291,25 @@ def task(request: Request, spec: str, ref: str) -> HTMLResponse:
|
|
| 262 |
path = f"/t/{enc(spec)}/{enc(ref)}"
|
| 263 |
cr, cld = crumbs(base, [("Environments", "/"), (env_name, f"/d/{enc(spec)}"), (title, "")])
|
| 264 |
body = (cr + f"<header class=\"tp-head\"><h1>{esc(title)}</h1><p class=\"lede\">{esc(desc)}</p></header>"
|
|
|
|
| 265 |
+ f'<p>Part of <a href="/d/{enc(spec)}">{esc(spec)}</a>.</p>')
|
|
|
|
|
|
|
|
|
|
|
|
|
| 266 |
ld = [{"@type": "CreativeWork", "name": title, "description": desc, "url": base + path, "learningResourceType": "RL environment task",
|
| 267 |
"isPartOf": {"@type": "Dataset", "name": env_name, "url": f"{base}/d/{enc(spec)}", "sameAs": f"https://huggingface.co/datasets/{spec}"},
|
| 268 |
"about": "reinforcement learning", **({"genre": row["category"]} if row and row.get("category") else {})}, cld]
|
| 269 |
return page(request, title=f"{title} Β· {env_name}", description=desc, path=path, body=body, jsonld=ld, image=f"/og{path}.png" if known else None,
|
| 270 |
-
noindex=not row, og_type="article")
|
| 271 |
|
| 272 |
|
| 273 |
def space(request: Request, spec: str) -> HTMLResponse:
|
| 274 |
base = base_url(request)
|
| 275 |
by, _ = listing()
|
| 276 |
-
r = by.
|
|
|
|
|
|
|
| 277 |
name = display_name(r, spec)
|
| 278 |
kind = kind_of(r, spec) if r else "environment Space"
|
| 279 |
desc = clip(f"{name}: an {kind} on Hugging Face. {clip((r or {}).get('brief') or '', 220) or 'See it live: its app, a playground with rewards, its tasks, and MCP for coding agents.'}", 300)
|
|
@@ -286,7 +322,49 @@ def space(request: Request, spec: str) -> HTMLResponse:
|
|
| 286 |
"offers": {"@type": "Offer", "price": "0", "priceCurrency": "USD"},
|
| 287 |
"author": {"@type": "Organization", "name": spec.split("/")[0], "url": f"https://huggingface.co/{spec.split('/')[0]}"},
|
| 288 |
"keywords": ["reinforcement learning", "RL environment", *(["OpenEnv"] if (r or {}).get("openenv") else []), *((r or {}).get("tags") or [])[:10]]}, cld]
|
| 289 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 290 |
|
| 291 |
|
| 292 |
def simple(request: Request, path: str) -> HTMLResponse:
|
|
@@ -310,7 +388,8 @@ def robots(request: Request) -> PlainTextResponse:
|
|
| 310 |
base = base_url(request)
|
| 311 |
return PlainTextResponse("\n".join([
|
| 312 |
"User-agent: *", "Allow: /", "Disallow: /api/", "Disallow: /mcp/", "Disallow: /capture/", "Disallow: /run/", "Disallow: /runs",
|
| 313 |
-
"
|
|
|
|
| 314 |
|
| 315 |
|
| 316 |
_sitemap: dict[str, Any] = {"at": 0.0, "pages": [], "tasks": []}
|
|
@@ -318,29 +397,45 @@ _sitemap: dict[str, Any] = {"at": 0.0, "pages": [], "tasks": []}
|
|
| 318 |
|
| 319 |
def _entries() -> tuple[list[tuple[str, str | None]], list[tuple[str, str | None]]]:
|
| 320 |
with _lock:
|
| 321 |
-
if time.time() - _sitemap["at"] <
|
| 322 |
return _sitemap["pages"], _sitemap["tasks"]
|
| 323 |
by, rows = listing()
|
|
|
|
| 324 |
pages: list[tuple[str, str | None]] = [("/", None), ("/community", None)]
|
| 325 |
-
#
|
| 326 |
pages += [(_href(r), (r.get("updated") or "")[:10] or None) for r in rows
|
| 327 |
-
if
|
| 328 |
-
if f"{MIMO}" not in by:
|
| 329 |
-
pages.append((f"/d/{enc(MIMO)}", None))
|
| 330 |
tasks: list[tuple[str, str | None]] = []
|
| 331 |
-
for spec in
|
| 332 |
try:
|
| 333 |
tasks += [(f"/t/{enc(spec)}/{enc(t['path'])}", None) for t in index_rows(spec)]
|
| 334 |
except Exception: # noqa: BLE001 - one index unreadable leaves the rest
|
| 335 |
continue
|
| 336 |
with _lock:
|
|
|
|
| 337 |
_sitemap.update(at=time.time(), pages=pages, tasks=tasks)
|
| 338 |
return pages, tasks
|
| 339 |
|
| 340 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 341 |
def _urlset(base: str, items: list[tuple[str, str | None]]) -> Response:
|
| 342 |
xml = ['<?xml version="1.0" encoding="UTF-8"?>', '<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
|
| 343 |
-
xml += [f"<url><loc>{esc(base + p)}</loc>{f'<lastmod>{esc(m)}</lastmod>' if m else ''}</url>" for p, m in items]
|
| 344 |
xml.append("</urlset>")
|
| 345 |
return Response("\n".join(xml), media_type="application/xml")
|
| 346 |
|
|
@@ -349,6 +444,8 @@ def sitemap_index(request: Request) -> Response:
|
|
| 349 |
base = base_url(request)
|
| 350 |
_, tasks = _entries()
|
| 351 |
parts = ["/sitemap-pages.xml", *[f"/sitemap-tasks-{i + 1}.xml" for i in range((len(tasks) + SITEMAP_CHUNK - 1) // SITEMAP_CHUNK)]]
|
|
|
|
|
|
|
| 352 |
xml = ['<?xml version="1.0" encoding="UTF-8"?>', '<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
|
| 353 |
xml += [f"<sitemap><loc>{esc(base + p)}</loc></sitemap>" for p in parts]
|
| 354 |
xml.append("</sitemapindex>")
|
|
@@ -368,6 +465,23 @@ def sitemap_tasks(request: Request, n: int) -> Response:
|
|
| 368 |
return _urlset(base_url(request), chunk)
|
| 369 |
|
| 370 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 371 |
# ββ preview images βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 372 |
# 1200Γ630, plain: the site's name, what the page is, its title and a line of facts. Drawn once per page and kept.
|
| 373 |
W, H = 1200, 630
|
|
@@ -377,27 +491,41 @@ W, H = 1200, 630
|
|
| 377 |
def card_png(kind: str, title: str, sub: str, facts: str) -> bytes:
|
| 378 |
from PIL import Image, ImageDraw, ImageFont
|
| 379 |
|
| 380 |
-
img = Image.new("RGB", (W, H), "#
|
| 381 |
d = ImageDraw.Draw(img)
|
| 382 |
font = lambda s: ImageFont.load_default(size=s) # noqa: E731 - Pillow's own font, so no system fonts are needed
|
| 383 |
-
d.rectangle([0, 0, W,
|
| 384 |
-
d.
|
| 385 |
-
d.text((
|
| 386 |
-
|
| 387 |
-
|
| 388 |
-
|
| 389 |
-
|
|
|
|
|
|
|
|
|
|
| 390 |
if sub:
|
| 391 |
-
d.text((
|
| 392 |
-
|
| 393 |
-
|
| 394 |
-
d.text((
|
| 395 |
out = io.BytesIO()
|
| 396 |
img.save(out, "PNG", optimize=True)
|
| 397 |
return out.getvalue()
|
| 398 |
|
| 399 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 400 |
def _wrap(d, text: str, font, width: int) -> list[str]:
|
|
|
|
|
|
|
| 401 |
words, lines, cur = str(text).split(), [], ""
|
| 402 |
for w in words:
|
| 403 |
nxt = f"{cur} {w}".strip()
|
|
@@ -415,9 +543,14 @@ def _wrap(d, text: str, font, width: int) -> list[str]:
|
|
| 415 |
return lines or [""]
|
| 416 |
|
| 417 |
|
| 418 |
-
def og_image(path: str) -> Response:
|
| 419 |
"""The preview image of a page, from its path (`/d/org/name`, `/t/org/name/ref`, `/s/org/name`, or the site's)."""
|
|
|
|
|
|
|
|
|
|
|
|
|
| 420 |
by, _ = listing()
|
|
|
|
| 421 |
parts = [p for p in path.strip("/").split("/") if p]
|
| 422 |
if parts and not (len(parts) >= 3 and parts[0] in ("d", "s", "t")):
|
| 423 |
return Response(status_code=404)
|
|
@@ -425,14 +558,30 @@ def og_image(path: str) -> Response:
|
|
| 425 |
if len(parts) >= 3 and parts[0] in ("d", "s", "t"):
|
| 426 |
spec = f"{parts[1]}/{parts[2]}"
|
| 427 |
r = by.get(spec if parts[0] != "s" else f"space:{spec}")
|
| 428 |
-
if r is None
|
| 429 |
return Response(status_code=404)
|
| 430 |
name = display_name(r, spec)
|
| 431 |
k = kind_of(r, spec)
|
| 432 |
if parts[0] == "t" and (r or spec == MIMO):
|
| 433 |
ref = "/".join(parts[3:])
|
| 434 |
row = next((t for t in index_rows(spec) if t["path"] == ref), None)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 435 |
kind, title, sub = f"A task in {name}", clip((row or {}).get("title") or ref, 140), spec
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 436 |
else:
|
| 437 |
n = ((r or {}).get("indexed") or {}).get("tasks") or (len(index_rows(spec)) if spec == MIMO else None)
|
| 438 |
kind, title, sub = k, name, spec
|
|
|
|
| 4 |
in its body, the page's content as plain HTML, for crawlers that don't run JavaScript (the app replaces it as it
|
| 5 |
starts). Plus robots.txt, a sitemap of every environment and every indexed task, and a preview image per page.
|
| 6 |
|
| 7 |
+
Catalog pages use public indexes. Row and Space task pages may make bounded anonymous
|
| 8 |
+
reads of the same withheld task views shown in the UI. Crawlers never start indexing,
|
| 9 |
+
wake a Space, run an episode or receive a visitor's credentials.
|
| 10 |
"""
|
| 11 |
|
| 12 |
from __future__ import annotations
|
|
|
|
| 19 |
import time
|
| 20 |
from functools import lru_cache
|
| 21 |
from typing import Any
|
| 22 |
+
from urllib.parse import quote, urlencode
|
| 23 |
|
| 24 |
from fastapi import Request
|
| 25 |
from fastapi.responses import HTMLResponse, PlainTextResponse, Response
|
| 26 |
|
| 27 |
+
from . import catalog, config, seo_tasks, snapshot, space_checks, spaces_live
|
| 28 |
|
| 29 |
SITE = "HF RL Explorer"
|
| 30 |
+
TAGLINE = "RL environments on the Hugging Face Hub"
|
| 31 |
+
DESCRIPTION = ("Explore reinforcement learning environments and tasks on the Hugging Face Hub. Browse OpenEnv, Harbor, "
|
| 32 |
+
"MiMo, NeMo Gym and Verifiers, inspect rewards, and run supported agent rollouts.")
|
| 33 |
KIND = {"harbor": "Harbor dataset", "verifiers": "Verifiers environment", "nemo-gym": "NeMo Gym dataset", "rows": "RL dataset",
|
| 34 |
"mimo": "MiMo RL release", "openenv": "OpenEnv Space", "space": "environment Space"}
|
| 35 |
MIMO = "XiaomiMiMo/MiMo-V2.6-RL-oss"
|
| 36 |
NOINDEX = re.compile(r"^/(run/|runs$|compare/)")
|
| 37 |
+
SITEMAP_CHUNK = 10_000
|
| 38 |
|
| 39 |
|
| 40 |
# ββ where we are βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
|
|
|
| 92 |
return by, rows
|
| 93 |
|
| 94 |
|
| 95 |
+
def public_rows(rows):
|
| 96 |
+
hidden = set(catalog.hidden())
|
| 97 |
+
return [r for r in rows if r["key"] not in hidden and not any(r.get(k) for k in ("private", "gated", "restricted"))]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
def discoverable(r):
|
| 101 |
+
return r.get("kind") == "dataset" or space_checks.browseable(space_checks.inventory().get(r["id"]))
|
| 102 |
+
|
| 103 |
+
|
| 104 |
def index_rows(spec: str) -> list[dict[str, Any]]:
|
| 105 |
"""An environment's tasks, if they are already known here (never built for a crawler): [{path, title, brief, category}]."""
|
| 106 |
if spec == MIMO:
|
|
|
|
| 109 |
p = catalog._index_path(spec)
|
| 110 |
mtime = p.stat().st_mtime
|
| 111 |
except (OSError, ValueError):
|
| 112 |
+
try:
|
| 113 |
+
return _snapshot_rows(spec, snapshot.get().name)
|
| 114 |
+
except snapshot.SnapshotError:
|
| 115 |
+
return []
|
| 116 |
return _harbor_rows(spec, mtime)
|
| 117 |
|
| 118 |
|
|
|
|
| 126 |
@lru_cache(maxsize=32)
|
| 127 |
def _harbor_rows(spec: str, mtime: float) -> list[dict[str, Any]]:
|
| 128 |
idx = catalog._read_index(spec)
|
| 129 |
+
if not idx or (idx.get("info") or {}).get("restricted"):
|
| 130 |
+
return []
|
| 131 |
return [{"path": t["path"], "title": t.get("title"), "brief": t.get("brief"), "category": t.get("category")} for t in (idx or {}).get("tasks") or []]
|
| 132 |
|
| 133 |
|
| 134 |
+
@lru_cache(maxsize=32)
|
| 135 |
+
def _snapshot_rows(spec: str, revision: str) -> list[dict[str, Any]]:
|
| 136 |
+
with snapshot.use() as (_, conn):
|
| 137 |
+
return [dict(r) for r in conn.execute("SELECT ref AS path, title, brief, category FROM tasks WHERE env = ? ORDER BY ref", (spec,))]
|
| 138 |
+
|
| 139 |
+
|
| 140 |
def display_name(row: dict[str, Any] | None, spec: str) -> str:
|
| 141 |
"""What to call an environment: its card's heading when that names it, else the repository name with the heading."""
|
| 142 |
repo = spec.split("/")[-1]
|
|
|
|
| 165 |
base = base_url(request)
|
| 166 |
url = base + path
|
| 167 |
full = title if SITE in title else f"{title} Β· {SITE}"
|
| 168 |
+
img = base + (image or "/social/rl-explorer.png")
|
| 169 |
head = "\n".join([
|
| 170 |
f'<link rel="canonical" href="{esc(url)}">',
|
| 171 |
+
f'<meta name="rlx-page-url" content="{esc(url)}">',
|
| 172 |
'<meta name="robots" content="noindex, follow">' if noindex else '<meta name="robots" content="index, follow, max-image-preview:large, max-snippet:-1">',
|
| 173 |
f'<meta property="og:site_name" content="{SITE}">',
|
| 174 |
f'<meta property="og:type" content="{og_type}">',
|
|
|
|
| 176 |
f'<meta property="og:description" content="{esc(description)}">',
|
| 177 |
f'<meta property="og:url" content="{esc(url)}">',
|
| 178 |
f'<meta property="og:image" content="{esc(img)}">',
|
| 179 |
+
'<meta property="og:image:type" content="image/png">',
|
| 180 |
+
'<meta property="og:locale" content="en_US">',
|
| 181 |
'<meta property="og:image:width" content="1200">',
|
| 182 |
'<meta property="og:image:height" content="630">',
|
| 183 |
f'<meta property="og:image:alt" content="{esc(full)}">',
|
|
|
|
| 185 |
f'<meta name="twitter:title" content="{esc(full)}">',
|
| 186 |
f'<meta name="twitter:description" content="{esc(description)}">',
|
| 187 |
f'<meta name="twitter:image" content="{esc(img)}">',
|
| 188 |
+
f'<meta name="twitter:image:alt" content="{esc(full)}">',
|
| 189 |
*[f'<script type="application/ld+json">{_ld(x)}</script>' for x in jsonld or []],
|
| 190 |
])
|
| 191 |
doc = template()
|
|
|
|
| 194 |
doc = doc.replace("</head>", f"{head}\n</head>", 1)
|
| 195 |
if body: # the page as plain HTML until the app starts (crawlers that don't run JavaScript read this)
|
| 196 |
doc = re.sub(r'<main id="view">(.*?)</main>', lambda m: f'<main id="view">{m.group(1)}<div class="wrap page ssr">{body}</div></main>', doc, count=1, flags=re.S)
|
| 197 |
+
return HTMLResponse(doc, status_code=status, headers={"Cache-Control": "no-cache", **({"Retry-After": "60"} if status == 503 else {})})
|
| 198 |
|
| 199 |
|
| 200 |
def _ld(x: dict[str, Any]) -> str:
|
|
|
|
| 215 |
def home(request: Request) -> HTMLResponse:
|
| 216 |
base = base_url(request)
|
| 217 |
by, rows = listing()
|
| 218 |
+
rows = [r for r in public_rows(rows) if discoverable(r)]
|
| 219 |
top = sorted(rows, key=lambda r: (r.get("trending") or 0) * 1e9 + (r.get("downloads") or 0), reverse=True)[:120]
|
| 220 |
+
desc = DESCRIPTION
|
|
|
|
|
|
|
|
|
|
|
|
|
| 221 |
body = (f"<header class=\"tp-head\"><h1>{TAGLINE}</h1><p class=\"lede\">{esc(desc)}</p></header>"
|
| 222 |
+ _env_list(top) + '<p><a href="/community">Community rollouts</a> Β· <a href="/d/XiaomiMiMo/MiMo-V2.6-RL-oss">MiMo-V2.6 RL</a></p>')
|
| 223 |
ld = [{"@type": "WebSite", "name": SITE, "alternateName": TAGLINE, "url": base + "/", "description": desc,
|
| 224 |
"potentialAction": {"@type": "SearchAction", "target": {"@type": "EntryPoint", "urlTemplate": base + "/?q={search_term_string}"},
|
| 225 |
"query-input": "required name=search_term_string"},
|
| 226 |
+
"publisher": {"@type": "Organization", "name": "FineEnvs", "url": "https://huggingface.co/FineEnvs"}},
|
| 227 |
{"@type": "CollectionPage", "name": TAGLINE, "url": base + "/", "about": "Reinforcement learning environments",
|
| 228 |
"mainEntity": {"@type": "ItemList", "numberOfItems": len(top), "itemListElement": [
|
| 229 |
{"@type": "ListItem", "position": i + 1, "url": base + _href(r), "name": r.get("heading") or r["id"]} for i, r in enumerate(top[:50])]}}]
|
|
|
|
| 242 |
def environment(request: Request, spec: str) -> HTMLResponse:
|
| 243 |
base = base_url(request)
|
| 244 |
by, _ = listing()
|
| 245 |
+
r = next((r for r in public_rows(list(by.values())) if r["key"] == spec), None)
|
| 246 |
name = display_name(r, spec)
|
| 247 |
kind = kind_of(r, spec)
|
| 248 |
+
tasks = index_rows(spec) if r else []
|
| 249 |
n = len(tasks) or ((r or {}).get("indexed") or {}).get("tasks")
|
| 250 |
brief = clip((r or {}).get("brief") or "", 300)
|
| 251 |
desc = clip(f"{name}: {kind} on Hugging Face{f' with {n:,} tasks' if n else ''}. "
|
|
|
|
| 264 |
"includedInDataCatalog": {"@type": "DataCatalog", "name": SITE, "url": base + "/"},
|
| 265 |
"distribution": [{"@type": "DataDownload", "encodingFormat": "application/octet-stream", "contentUrl": f"https://huggingface.co/datasets/{spec}"}]},
|
| 266 |
cld]
|
| 267 |
+
return page(request, title=f"{name} Β· {kind}", description=desc, path=path, body=body, jsonld=ld, image=f"/og{path}.png" if r else None, noindex=not r)
|
| 268 |
|
| 269 |
|
| 270 |
def task(request: Request, spec: str, ref: str) -> HTMLResponse:
|
| 271 |
base = base_url(request)
|
| 272 |
by, _ = listing()
|
| 273 |
+
r = next((r for r in public_rows(list(by.values())) if r["key"] == spec), None)
|
| 274 |
env_name = display_name(r, spec)
|
| 275 |
+
known = r is not None
|
| 276 |
row = next((t for t in index_rows(spec) if t["path"] == ref), None) if known else None
|
| 277 |
+
status = 200
|
| 278 |
+
if known and not row and re.fullmatch(r"(?:[^/]+/){1,2}\d{1,10}", ref):
|
| 279 |
+
try:
|
| 280 |
+
row = seo_tasks.row(spec, ref)
|
| 281 |
+
except seo_tasks.Unavailable as exc:
|
| 282 |
+
status = exc.status
|
| 283 |
+
elif known and not row:
|
| 284 |
+
status = 404
|
| 285 |
split, _, n = ref.rpartition("/")
|
| 286 |
# a row of a dataset read as rows (no index here): "Row 3 (train)" rather than a bare "3"
|
| 287 |
fallback = f"Row {n} ({split})" if split and n.isdigit() and "/" not in split else ref.rsplit("/", 1)[-1]
|
|
|
|
| 291 |
path = f"/t/{enc(spec)}/{enc(ref)}"
|
| 292 |
cr, cld = crumbs(base, [("Environments", "/"), (env_name, f"/d/{enc(spec)}"), (title, "")])
|
| 293 |
body = (cr + f"<header class=\"tp-head\"><h1>{esc(title)}</h1><p class=\"lede\">{esc(desc)}</p></header>"
|
| 294 |
+
+ (f'<section><h2>The task</h2><p class="seo-prompt">{esc(row.get("brief") or "")}</p></section>' if row else "")
|
| 295 |
+ f'<p>Part of <a href="/d/{enc(spec)}">{esc(spec)}</a>.</p>')
|
| 296 |
+
if row and type(row.get("total")) is int and n.isdigit():
|
| 297 |
+
body += '<nav aria-label="Other tasks">' + " Β· ".join(
|
| 298 |
+
f'<a href="/t/{enc(spec)}/{enc(split)}/{i}">{label}</a>'
|
| 299 |
+
for i, label in ((int(n) - 1, "Previous task"), (int(n) + 1, "Next task")) if 0 <= i < row["total"]) + '</nav>'
|
| 300 |
ld = [{"@type": "CreativeWork", "name": title, "description": desc, "url": base + path, "learningResourceType": "RL environment task",
|
| 301 |
"isPartOf": {"@type": "Dataset", "name": env_name, "url": f"{base}/d/{enc(spec)}", "sameAs": f"https://huggingface.co/datasets/{spec}"},
|
| 302 |
"about": "reinforcement learning", **({"genre": row["category"]} if row and row.get("category") else {})}, cld]
|
| 303 |
return page(request, title=f"{title} Β· {env_name}", description=desc, path=path, body=body, jsonld=ld, image=f"/og{path}.png" if known else None,
|
| 304 |
+
noindex=not row and status != 503, og_type="article", status=status)
|
| 305 |
|
| 306 |
|
| 307 |
def space(request: Request, spec: str) -> HTMLResponse:
|
| 308 |
base = base_url(request)
|
| 309 |
by, _ = listing()
|
| 310 |
+
r = next((r for r in public_rows(list(by.values())) if r["key"] == f"space:{spec}"), None)
|
| 311 |
+
if "task" in request.query_params:
|
| 312 |
+
return space_task(request, spec, r)
|
| 313 |
name = display_name(r, spec)
|
| 314 |
kind = kind_of(r, spec) if r else "environment Space"
|
| 315 |
desc = clip(f"{name}: an {kind} on Hugging Face. {clip((r or {}).get('brief') or '', 220) or 'See it live: its app, a playground with rewards, its tasks, and MCP for coding agents.'}", 300)
|
|
|
|
| 322 |
"offers": {"@type": "Offer", "price": "0", "priceCurrency": "USD"},
|
| 323 |
"author": {"@type": "Organization", "name": spec.split("/")[0], "url": f"https://huggingface.co/{spec.split('/')[0]}"},
|
| 324 |
"keywords": ["reinforcement learning", "RL environment", *(["OpenEnv"] if (r or {}).get("openenv") else []), *((r or {}).get("tags") or [])[:10]]}, cld]
|
| 325 |
+
if r and discoverable(r):
|
| 326 |
+
ranges = seo_tasks.ranges((spaces_live.last_seen(spec) or {}).get("task_api"))
|
| 327 |
+
body += "".join(f'<section><h2>{esc(split)} tasks</h2><ul>' + "".join(
|
| 328 |
+
f'<li><a href="{esc(seo_tasks.space_path(spec, env, split, i))}">Task {i + 1}</a></li>' for i in range(min(n, 12)))
|
| 329 |
+
+ "</ul></section>" for env, split, n in ranges)
|
| 330 |
+
return page(request, title=f"{name} Β· {kind}", description=desc, path=path, body=body, jsonld=ld, image=f"/og{path}.png" if r else None, noindex=not r or not discoverable(r))
|
| 331 |
+
|
| 332 |
+
|
| 333 |
+
def space_task(request: Request, spec: str, environment: dict | None) -> HTMLResponse:
|
| 334 |
+
q = request.query_params
|
| 335 |
+
env, split, raw = q.get("env", ""), q.get("split", ""), q.get("task", "")
|
| 336 |
+
if not environment or not re.fullmatch(r"\d{1,10}", raw) or not env or not split or len(env) > 80 or len(split) > 200:
|
| 337 |
+
return not_found(request)
|
| 338 |
+
index = int(raw)
|
| 339 |
+
path = seo_tasks.space_path(spec, env, split, index)
|
| 340 |
+
status, row = 200, None
|
| 341 |
+
try:
|
| 342 |
+
row = seo_tasks.space(spec, env, split, index)
|
| 343 |
+
except seo_tasks.Unavailable as exc:
|
| 344 |
+
status = exc.status
|
| 345 |
+
except Exception:
|
| 346 |
+
status = 503
|
| 347 |
+
if status == 404:
|
| 348 |
+
return not_found(request)
|
| 349 |
+
title = clip((row or {}).get("title") or f"Task {index + 1}", 110)
|
| 350 |
+
desc = clip(f"{title}: {split} task in {spec}, an RL environment on the Hugging Face Hub. {(row or {}).get('brief') or ''}", 300)
|
| 351 |
+
cr, cld = crumbs(base_url(request), [("Environments", "/"), (spec, f"/s/{enc(spec)}"), (title, "")])
|
| 352 |
+
body = cr + f'<header class="tp-head"><h1>{esc(title)}</h1><p class="lede">{esc(desc)}</p></header>'
|
| 353 |
+
if row:
|
| 354 |
+
body += f'<section><h2>The task</h2><p class="seo-prompt">{esc(row["brief"])}</p></section>'
|
| 355 |
+
if row.get("fields"):
|
| 356 |
+
body += '<h2>Task details</h2><dl>' + "".join(
|
| 357 |
+
f'<dt>{esc(k.replace("_", " "))}</dt><dd>{esc(str(v)[:1000])}</dd>' for k, v in row["fields"].items()) + '</dl>'
|
| 358 |
+
body += '<nav aria-label="Other tasks">' + " Β· ".join(
|
| 359 |
+
f'<a href="{esc(seo_tasks.space_path(spec, env, split, i))}">{label}</a>'
|
| 360 |
+
for i, label in ((index - 1, "Previous task"), (index + 1, "Next task")) if 0 <= i < row["total"]) + '</nav>'
|
| 361 |
+
else:
|
| 362 |
+
body += '<p>The task server is temporarily unavailable. Please try again shortly.</p>'
|
| 363 |
+
ld = [{"@type": "CreativeWork", "name": title, "description": desc, "url": base_url(request) + path,
|
| 364 |
+
"isPartOf": {"@type": "SoftwareApplication", "name": spec, "url": base_url(request) + f"/s/{enc(spec)}"}}, cld]
|
| 365 |
+
return page(request, title=f"{title} Β· {spec}", description=desc, path=path, body=body, jsonld=ld,
|
| 366 |
+
image=f"/og/s/{enc(spec)}.png?" + urlencode({"env": env, "split": split, "task": index}),
|
| 367 |
+
og_type="article", status=status)
|
| 368 |
|
| 369 |
|
| 370 |
def simple(request: Request, path: str) -> HTMLResponse:
|
|
|
|
| 388 |
base = base_url(request)
|
| 389 |
return PlainTextResponse("\n".join([
|
| 390 |
"User-agent: *", "Allow: /", "Disallow: /api/", "Disallow: /mcp/", "Disallow: /capture/", "Disallow: /run/", "Disallow: /runs",
|
| 391 |
+
"Allow: /api/env/", "Allow: /api/spaces/", "Allow: /api/search", "Allow: /api/environments",
|
| 392 |
+
"Disallow: /api/environments/mine", "Disallow: /compare/", "Disallow: /login", "Disallow: /logout", "", f"Sitemap: {base}/sitemap.xml", ""]))
|
| 393 |
|
| 394 |
|
| 395 |
_sitemap: dict[str, Any] = {"at": 0.0, "pages": [], "tasks": []}
|
|
|
|
| 397 |
|
| 398 |
def _entries() -> tuple[list[tuple[str, str | None]], list[tuple[str, str | None]]]:
|
| 399 |
with _lock:
|
| 400 |
+
if time.time() - _sitemap["at"] < 300 and _sitemap["pages"]:
|
| 401 |
return _sitemap["pages"], _sitemap["tasks"]
|
| 402 |
by, rows = listing()
|
| 403 |
+
rows = public_rows(rows)
|
| 404 |
pages: list[tuple[str, str | None]] = [("/", None), ("/community", None)]
|
| 405 |
+
# Every public dataset and only Spaces with evidence of a supported API.
|
| 406 |
pages += [(_href(r), (r.get("updated") or "")[:10] or None) for r in rows
|
| 407 |
+
if discoverable(r)]
|
|
|
|
|
|
|
| 408 |
tasks: list[tuple[str, str | None]] = []
|
| 409 |
+
for spec in dict.fromkeys(r["id"] for r in rows if r["kind"] == "dataset"):
|
| 410 |
try:
|
| 411 |
tasks += [(f"/t/{enc(spec)}/{enc(t['path'])}", None) for t in index_rows(spec)]
|
| 412 |
except Exception: # noqa: BLE001 - one index unreadable leaves the rest
|
| 413 |
continue
|
| 414 |
with _lock:
|
| 415 |
+
tasks = list(dict.fromkeys(tasks))
|
| 416 |
_sitemap.update(at=time.time(), pages=pages, tasks=tasks)
|
| 417 |
return pages, tasks
|
| 418 |
|
| 419 |
|
| 420 |
+
def _space_ranges():
|
| 421 |
+
"""Represent task coordinates as ranges; do not allocate millions of URLs or fetch any tasks."""
|
| 422 |
+
records = space_checks.inventory()
|
| 423 |
+
_, rows = listing()
|
| 424 |
+
out = []
|
| 425 |
+
for r in sorted(public_rows(rows), key=lambda r: r["key"]):
|
| 426 |
+
if r.get("kind") != "space" or not space_checks.browseable(records.get(r["id"])):
|
| 427 |
+
continue
|
| 428 |
+
rec = records[r["id"]]
|
| 429 |
+
ranges = rec.get("task_splits") or seo_tasks.ranges((spaces_live.last_seen(r["id"]) or {}).get("task_api"))
|
| 430 |
+
for env, split, n in ranges:
|
| 431 |
+
if type(n) is int and 0 < n <= 1_000_000_000:
|
| 432 |
+
out.append((r["id"], env, split, n))
|
| 433 |
+
return out
|
| 434 |
+
|
| 435 |
+
|
| 436 |
def _urlset(base: str, items: list[tuple[str, str | None]]) -> Response:
|
| 437 |
xml = ['<?xml version="1.0" encoding="UTF-8"?>', '<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
|
| 438 |
+
xml += [f"<url><loc>{esc(base + p)}</loc>{f'<lastmod>{esc(m)}</lastmod>' if m else ''}</url>" for p, m in items if len((base + p).encode()) <= 2048]
|
| 439 |
xml.append("</urlset>")
|
| 440 |
return Response("\n".join(xml), media_type="application/xml")
|
| 441 |
|
|
|
|
| 444 |
base = base_url(request)
|
| 445 |
_, tasks = _entries()
|
| 446 |
parts = ["/sitemap-pages.xml", *[f"/sitemap-tasks-{i + 1}.xml" for i in range((len(tasks) + SITEMAP_CHUNK - 1) // SITEMAP_CHUNK)]]
|
| 447 |
+
total = sum(r[3] for r in _space_ranges())
|
| 448 |
+
parts += [f"/sitemap-space-tasks-{i + 1}.xml" for i in range(min(49000, (total + SITEMAP_CHUNK - 1) // SITEMAP_CHUNK))]
|
| 449 |
xml = ['<?xml version="1.0" encoding="UTF-8"?>', '<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
|
| 450 |
xml += [f"<sitemap><loc>{esc(base + p)}</loc></sitemap>" for p in parts]
|
| 451 |
xml.append("</sitemapindex>")
|
|
|
|
| 465 |
return _urlset(base_url(request), chunk)
|
| 466 |
|
| 467 |
|
| 468 |
+
def sitemap_space_tasks(request: Request, n: int) -> Response:
|
| 469 |
+
if not 1 <= n <= 49000:
|
| 470 |
+
return Response("not found", status_code=404)
|
| 471 |
+
offset, remaining, urls = (n - 1) * SITEMAP_CHUNK, SITEMAP_CHUNK, []
|
| 472 |
+
for spec, env, split, count in _space_ranges():
|
| 473 |
+
if offset >= count:
|
| 474 |
+
offset -= count
|
| 475 |
+
continue
|
| 476 |
+
stop = min(count, offset + remaining)
|
| 477 |
+
urls.extend((seo_tasks.space_path(spec, env, split, i), None) for i in range(offset, stop))
|
| 478 |
+
remaining -= stop - offset
|
| 479 |
+
offset = 0
|
| 480 |
+
if not remaining:
|
| 481 |
+
break
|
| 482 |
+
return _urlset(base_url(request), urls) if urls else Response("not found", status_code=404)
|
| 483 |
+
|
| 484 |
+
|
| 485 |
# ββ preview images βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
| 486 |
# 1200Γ630, plain: the site's name, what the page is, its title and a line of facts. Drawn once per page and kept.
|
| 487 |
W, H = 1200, 630
|
|
|
|
| 491 |
def card_png(kind: str, title: str, sub: str, facts: str) -> bytes:
|
| 492 |
from PIL import Image, ImageDraw, ImageFont
|
| 493 |
|
| 494 |
+
img = Image.new("RGB", (W, H), "#fffdf7")
|
| 495 |
d = ImageDraw.Draw(img)
|
| 496 |
font = lambda s: ImageFont.load_default(size=s) # noqa: E731 - Pillow's own font, so no system fonts are needed
|
| 497 |
+
d.rectangle([0, 0, W, 12], fill="#ffcd36")
|
| 498 |
+
d.rounded_rectangle([56, 48, 198, 90], radius=10, fill="#ffdc62")
|
| 499 |
+
d.text((73, 58), "FINEENVS", font=font(23), fill="#352c0e")
|
| 500 |
+
d.text((220, 57), SITE, font=font(26), fill="#54504a")
|
| 501 |
+
d.text((56, 142), _ellipsize(d, kind.upper(), font(23), 1088), font=font(23), fill="#827256")
|
| 502 |
+
lines = _wrap(d, title, font(65), 1088)
|
| 503 |
+
size = 65 if len(lines) <= 3 else 54
|
| 504 |
+
lines = _wrap(d, title, font(size), 1088)[:3]
|
| 505 |
+
for i, line in enumerate(lines):
|
| 506 |
+
d.text((56, 194 + i * 77), _ellipsize(d, line, font(size), 1088), font=font(size), fill="#1c1b19")
|
| 507 |
if sub:
|
| 508 |
+
d.text((56, 452), _ellipsize(d, sub, font(27), 1088), font=font(27), fill="#665e51")
|
| 509 |
+
d.line([56, 516, 1144, 516], fill="#e5dfd1", width=2)
|
| 510 |
+
d.text((56, 550), _ellipsize(d, facts or "Discover tasks. Inspect rewards. Run agents.", font(24), 680), font=font(24), fill="#524d42")
|
| 511 |
+
d.text((1144, 550), "Hugging Face Hub", font=font(24), fill="#827256", anchor="ra")
|
| 512 |
out = io.BytesIO()
|
| 513 |
img.save(out, "PNG", optimize=True)
|
| 514 |
return out.getvalue()
|
| 515 |
|
| 516 |
|
| 517 |
+
def _ellipsize(d, text, font, width):
|
| 518 |
+
text = str(text)
|
| 519 |
+
if d.textlength(text, font=font) <= width:
|
| 520 |
+
return text
|
| 521 |
+
while text and d.textlength(text + "β¦", font=font) > width:
|
| 522 |
+
text = text[:-1]
|
| 523 |
+
return text + "β¦"
|
| 524 |
+
|
| 525 |
+
|
| 526 |
def _wrap(d, text: str, font, width: int) -> list[str]:
|
| 527 |
+
if "\n" in text:
|
| 528 |
+
return [line for part in text.splitlines() for line in _wrap(d, part, font, width)]
|
| 529 |
words, lines, cur = str(text).split(), [], ""
|
| 530 |
for w in words:
|
| 531 |
nxt = f"{cur} {w}".strip()
|
|
|
|
| 543 |
return lines or [""]
|
| 544 |
|
| 545 |
|
| 546 |
+
def og_image(path: str, query=None) -> Response:
|
| 547 |
"""The preview image of a page, from its path (`/d/org/name`, `/t/org/name/ref`, `/s/org/name`, or the site's)."""
|
| 548 |
+
if path == "/":
|
| 549 |
+
return Response(card_png("Reinforcement learning", "Explore RL environments\non the Hugging Face Hub",
|
| 550 |
+
"OpenEnv Β· Harbor Β· MiMo Β· NeMo Gym Β· Verifiers", ""),
|
| 551 |
+
media_type="image/png", headers={"Cache-Control": "public, max-age=86400"})
|
| 552 |
by, _ = listing()
|
| 553 |
+
by = {r["key"]: r for r in public_rows(list(by.values()))}
|
| 554 |
parts = [p for p in path.strip("/").split("/") if p]
|
| 555 |
if parts and not (len(parts) >= 3 and parts[0] in ("d", "s", "t")):
|
| 556 |
return Response(status_code=404)
|
|
|
|
| 558 |
if len(parts) >= 3 and parts[0] in ("d", "s", "t"):
|
| 559 |
spec = f"{parts[1]}/{parts[2]}"
|
| 560 |
r = by.get(spec if parts[0] != "s" else f"space:{spec}")
|
| 561 |
+
if r is None:
|
| 562 |
return Response(status_code=404)
|
| 563 |
name = display_name(r, spec)
|
| 564 |
k = kind_of(r, spec)
|
| 565 |
if parts[0] == "t" and (r or spec == MIMO):
|
| 566 |
ref = "/".join(parts[3:])
|
| 567 |
row = next((t for t in index_rows(spec) if t["path"] == ref), None)
|
| 568 |
+
if not row and re.fullmatch(r"(?:[^/]+/){1,2}\d{1,10}", ref):
|
| 569 |
+
try:
|
| 570 |
+
row = seo_tasks.row(spec, ref)
|
| 571 |
+
except seo_tasks.Unavailable as exc:
|
| 572 |
+
return Response(status_code=exc.status)
|
| 573 |
+
if not row:
|
| 574 |
+
return Response(status_code=404)
|
| 575 |
kind, title, sub = f"A task in {name}", clip((row or {}).get("title") or ref, 140), spec
|
| 576 |
+
elif parts[0] == "s" and query and "task" in query:
|
| 577 |
+
raw, env, split = query.get("task", ""), query.get("env", ""), query.get("split", "")
|
| 578 |
+
if not re.fullmatch(r"\d{1,10}", raw) or not env or not split or len(env) > 80 or len(split) > 200:
|
| 579 |
+
return Response(status_code=404)
|
| 580 |
+
try:
|
| 581 |
+
row = seo_tasks.space(spec, env, split, int(raw))
|
| 582 |
+
except Exception:
|
| 583 |
+
return Response(status_code=503, headers={"Retry-After": "60"})
|
| 584 |
+
kind, title, sub = f"{split} task", clip(row["title"], 140), spec
|
| 585 |
else:
|
| 586 |
n = ((r or {}).get("indexed") or {}).get("tasks") or (len(index_rows(spec)) if spec == MIMO else None)
|
| 587 |
kind, title, sub = k, name, spec
|
app/seo_tasks.py
ADDED
|
@@ -0,0 +1,115 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Anonymous, bounded task reads for server-rendered public task pages.
|
| 2 |
+
|
| 3 |
+
No episode, index build, credentials or model call. Only advertised Task API
|
| 4 |
+
coordinates and public row datasets may be read. Answers use the UI's withholding.
|
| 5 |
+
"""
|
| 6 |
+
from __future__ import annotations
|
| 7 |
+
|
| 8 |
+
import re
|
| 9 |
+
import threading
|
| 10 |
+
from urllib.parse import urlencode
|
| 11 |
+
|
| 12 |
+
from . import catalog, spaces_live
|
| 13 |
+
|
| 14 |
+
_reads = threading.BoundedSemaphore(4)
|
| 15 |
+
|
| 16 |
+
|
| 17 |
+
class Unavailable(Exception):
|
| 18 |
+
def __init__(self, status=503):
|
| 19 |
+
self.status = status
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
def read(key, fn):
|
| 23 |
+
def fetch():
|
| 24 |
+
if not _reads.acquire(blocking=False):
|
| 25 |
+
raise Unavailable()
|
| 26 |
+
try:
|
| 27 |
+
return fn()
|
| 28 |
+
finally:
|
| 29 |
+
_reads.release()
|
| 30 |
+
return catalog._cached(("seo-task", *key), 600, fetch)
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def space_path(spec, env, split, index):
|
| 34 |
+
return f"/s/{spec}?" + urlencode({"env": env, "split": split, "task": index})
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def ranges(task_api):
|
| 38 |
+
if not isinstance(task_api, dict):
|
| 39 |
+
return []
|
| 40 |
+
result, seen = [], set()
|
| 41 |
+
for e in (task_api.get("environments") or [task_api])[:16]:
|
| 42 |
+
name = e.get("env")
|
| 43 |
+
if not isinstance(name, str) or not re.fullmatch(r"[\w.-]{1,80}", name) or name in (".", ".."):
|
| 44 |
+
continue
|
| 45 |
+
for s in e.get("splits", [])[:40]:
|
| 46 |
+
split, n = s.get("name"), s.get("num_tasks")
|
| 47 |
+
if not isinstance(split, str) or not 0 < len(split) <= 200 or type(n) is not int or not 0 < n <= 1_000_000_000:
|
| 48 |
+
continue
|
| 49 |
+
if (name, split) not in seen:
|
| 50 |
+
result.append([name, split, n])
|
| 51 |
+
seen.add((name, split))
|
| 52 |
+
return result
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
def space(spec, env, split, index):
|
| 56 |
+
"""Validate against the last public probe before making a read-only task request."""
|
| 57 |
+
known = ranges((spaces_live.last_seen(spec) or {}).get("task_api"))
|
| 58 |
+
if not known:
|
| 59 |
+
known = ranges(spaces_live.probe(spec).get("task_api"))
|
| 60 |
+
match = next((r for r in known if r[0] == env and r[1] == split), None)
|
| 61 |
+
if not match or not 0 <= index < match[2]:
|
| 62 |
+
raise Unavailable(404)
|
| 63 |
+
|
| 64 |
+
def fetch():
|
| 65 |
+
try:
|
| 66 |
+
raw = spaces_live.task(spec, split, index, env).get("task")
|
| 67 |
+
except spaces_live.SpaceError as exc:
|
| 68 |
+
raise Unavailable(404 if exc.status == 404 else 503) from None
|
| 69 |
+
if not isinstance(raw, dict):
|
| 70 |
+
raise Unavailable(404)
|
| 71 |
+
# Explicit text fields only; do not serialize arbitrary task metadata.
|
| 72 |
+
title = next((raw[k] for k in ("task_name", "title", "task_id", "id") if isinstance(raw.get(k), str)), f"Task {index + 1}")
|
| 73 |
+
prompt = next((raw[k] for k in ("prompt", "instruction", "description", "question") if isinstance(raw.get(k), str)), "")
|
| 74 |
+
fields = {k: raw[k] for k in ("category", "difficulty", "language", "language_name", "family", "mime", "duration_seconds",
|
| 75 |
+
"sampling_rate", "n_frames", "provider", "sequence_id", "offline_ready", "media_ready")
|
| 76 |
+
if isinstance(raw.get(k), (str, int, float, bool))}
|
| 77 |
+
return {"title": title[:400], "brief": prompt[:16000], "total": match[2], "fields": fields}
|
| 78 |
+
return read(("space", spec, env, split, index), fetch)
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def row(spec, ref):
|
| 82 |
+
from .envs import rows
|
| 83 |
+
|
| 84 |
+
try:
|
| 85 |
+
config, split, index = rows.parse_ref(ref)
|
| 86 |
+
except (ValueError, LookupError):
|
| 87 |
+
raise Unavailable(404) from None
|
| 88 |
+
if index < 0:
|
| 89 |
+
raise Unavailable(404)
|
| 90 |
+
|
| 91 |
+
def fetch():
|
| 92 |
+
try:
|
| 93 |
+
view = rows.task(spec, config, split, index, token=None)
|
| 94 |
+
except (LookupError, PermissionError):
|
| 95 |
+
raise Unavailable(404) from None
|
| 96 |
+
except Exception:
|
| 97 |
+
raise Unavailable() from None
|
| 98 |
+
if view.get("restricted"):
|
| 99 |
+
raise Unavailable(404)
|
| 100 |
+
texts = []
|
| 101 |
+
for section in view.get("sections", []):
|
| 102 |
+
if section.get("id") not in ("task", "prompt", "messages", "instruction"):
|
| 103 |
+
continue
|
| 104 |
+
body = section.get("body")
|
| 105 |
+
if isinstance(body, str):
|
| 106 |
+
texts.append(body)
|
| 107 |
+
elif section.get("kind") == "messages" and isinstance(body, list):
|
| 108 |
+
texts.extend(m["content"] for m in body if isinstance(m, dict)
|
| 109 |
+
and m.get("role") in ("system", "user") and isinstance(m.get("content"), str))
|
| 110 |
+
elif section.get("kind") == "blocks" and isinstance(body, list):
|
| 111 |
+
texts.extend(b["text"] for b in body if isinstance(b, dict)
|
| 112 |
+
and b.get("type") in ("markdown", "custom", "note") and isinstance(b.get("text"), str))
|
| 113 |
+
return {"title": str(view.get("title") or ref)[:400], "brief": "\n\n".join(texts)[:16000],
|
| 114 |
+
"path": ref, "total": view.get("total")}
|
| 115 |
+
return read(("row", spec, ref), fetch)
|
app/space_checks.py
CHANGED
|
@@ -219,6 +219,17 @@ def assess(api, health, metadata, schema, tools):
|
|
| 219 |
"failed": [k for k, ok in checks.items() if not ok], "tools": len(tools or [])}
|
| 220 |
|
| 221 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 222 |
def check(spec):
|
| 223 |
from . import spaces_live as live
|
| 224 |
rec = {"schema": 1, "id": catalog.check_spec(spec), "checked_at": time.time(), "status": "Check unavailable"}
|
|
@@ -249,6 +260,7 @@ def check(spec):
|
|
| 249 |
from . import space_tasks
|
| 250 |
task_api = space_tasks.discover(hub, paths)
|
| 251 |
rec["task_catalog"] = space_tasks.summary(task_api)
|
|
|
|
| 252 |
if rec["status"] == PASS:
|
| 253 |
rec["version"] = declared_version(spec, hub.get("tags"))
|
| 254 |
elif rec["version"]["value"] == "Unknown":
|
|
@@ -273,6 +285,7 @@ def observe(spec, info):
|
|
| 273 |
rec["interface"] = environment_interface({"info": {"version": info.get("openapi_version")}, "paths": info.get("endpoint_methods")})
|
| 274 |
from . import space_tasks
|
| 275 |
rec["task_catalog"] = space_tasks.summary(info.get("task_api"))
|
|
|
|
| 276 |
info["openenv_verified"] = verified(rec)
|
| 277 |
save(rec)
|
| 278 |
except Exception:
|
|
|
|
| 219 |
"failed": [k for k, ok in checks.items() if not ok], "tools": len(tools or [])}
|
| 220 |
|
| 221 |
|
| 222 |
+
def _task_splits(task_api):
|
| 223 |
+
from .seo_tasks import ranges
|
| 224 |
+
|
| 225 |
+
out = []
|
| 226 |
+
for row in ranges(task_api):
|
| 227 |
+
if len(json.dumps(out + [row]).encode()) > 12000:
|
| 228 |
+
break
|
| 229 |
+
out.append(row)
|
| 230 |
+
return out
|
| 231 |
+
|
| 232 |
+
|
| 233 |
def check(spec):
|
| 234 |
from . import spaces_live as live
|
| 235 |
rec = {"schema": 1, "id": catalog.check_spec(spec), "checked_at": time.time(), "status": "Check unavailable"}
|
|
|
|
| 260 |
from . import space_tasks
|
| 261 |
task_api = space_tasks.discover(hub, paths)
|
| 262 |
rec["task_catalog"] = space_tasks.summary(task_api)
|
| 263 |
+
rec["task_splits"] = _task_splits(task_api)
|
| 264 |
if rec["status"] == PASS:
|
| 265 |
rec["version"] = declared_version(spec, hub.get("tags"))
|
| 266 |
elif rec["version"]["value"] == "Unknown":
|
|
|
|
| 285 |
rec["interface"] = environment_interface({"info": {"version": info.get("openapi_version")}, "paths": info.get("endpoint_methods")})
|
| 286 |
from . import space_tasks
|
| 287 |
rec["task_catalog"] = space_tasks.summary(info.get("task_api"))
|
| 288 |
+
rec["task_splits"] = _task_splits(info.get("task_api"))
|
| 289 |
info["openenv_verified"] = verified(rec)
|
| 290 |
save(rec)
|
| 291 |
except Exception:
|
web/index.html
CHANGED
|
@@ -3,7 +3,7 @@
|
|
| 3 |
<head>
|
| 4 |
<meta charset="utf-8">
|
| 5 |
<meta name="viewport" content="width=device-width, initial-scale=1, viewport-fit=cover">
|
| 6 |
-
<title>HF RL Explorer:
|
| 7 |
<link rel="icon" href="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" type="image/svg+xml">
|
| 8 |
<meta name="description" content="Explore reinforcement learning environments on Hugging Face across OpenEnv, Harbor, Verifiers, NeMo Gym and more: what each task asks, how it's graded, what it runs in, and run an agent on it.">
|
| 9 |
<meta name="application-name" content="HF RL Explorer">
|
|
|
|
| 3 |
<head>
|
| 4 |
<meta charset="utf-8">
|
| 5 |
<meta name="viewport" content="width=device-width, initial-scale=1, viewport-fit=cover">
|
| 6 |
+
<title>HF RL Explorer: RL environments on the Hugging Face Hub</title>
|
| 7 |
<link rel="icon" href="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" type="image/svg+xml">
|
| 8 |
<meta name="description" content="Explore reinforcement learning environments on Hugging Face across OpenEnv, Harbor, Verifiers, NeMo Gym and more: what each task asks, how it's graded, what it runs in, and run an agent on it.">
|
| 9 |
<meta name="application-name" content="HF RL Explorer">
|
web/js/seo-meta.js
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
export const SITE = "HF RL Explorer";
|
| 2 |
+
export const DESCRIPTION = "Explore reinforcement learning environments and tasks on the Hugging Face Hub. Browse OpenEnv, Harbor, MiMo, NeMo Gym and Verifiers, inspect rewards, and run supported agent rollouts.";
|
| 3 |
+
|
| 4 |
+
// A Space task is a distinct resource. Ignore filters and tracking parameters,
|
| 5 |
+
// but preserve its environment, split and numeric task index in a stable order.
|
| 6 |
+
export function canonicalPath(url) {
|
| 7 |
+
const u = new URL(url), q = u.searchParams;
|
| 8 |
+
if (/^\/s\/[^/]+\/[^/]+$/.test(u.pathname) && /^\d{1,10}$/.test(q.get("task") || "") && q.get("env") && q.get("split")) {
|
| 9 |
+
return u.pathname + "?" + new URLSearchParams({ env: q.get("env"), split: q.get("split"), task: String(Number(q.get("task"))) });
|
| 10 |
+
}
|
| 11 |
+
return u.pathname;
|
| 12 |
+
}
|
| 13 |
+
|
| 14 |
+
export function imagePath(url) {
|
| 15 |
+
const path = canonicalPath(url), [name, query] = path.split("?");
|
| 16 |
+
return /^\/(d|t|s)\//.test(name) ? `/og${name}.png${query ? `?${query}` : ""}` : "/social/rl-explorer.png";
|
| 17 |
+
}
|
web/js/space.js
CHANGED
|
@@ -782,14 +782,17 @@ function render(me) {
|
|
| 782 |
if (!s) return;
|
| 783 |
const [org, name] = spec.split("/");
|
| 784 |
const sections = sectionsOf(me);
|
|
|
|
|
|
|
| 785 |
const kind = me.L?.openenv_verified ? "OpenEnv Space" : s.declared_openenv || s.openenv || s.manifest ? "Unverified Space" : s.framework === "ors" ? "ORS Space" : "environment Space"; // as the server-rendered title says
|
| 786 |
-
setMeta({ title: `${s.heading || name} Β· ${kind}`, description: `${spec}: an RL environment Space on Hugging Face. ${describe(me)} See it live: its app, a playground with rewards, its tasks, and MCP for coding agents.` });
|
| 787 |
me.viewer?.destroy();
|
| 788 |
me.viewer = null;
|
| 789 |
el.innerHTML = `<div class="wrap page${me.rendered ? "" : " fade-in"}">
|
| 790 |
<nav class="crumbs" aria-label="Breadcrumb"><a href="/">Environments</a>${icon("chevronRight", 13)}<a href="/?k=space">Spaces</a>${icon("chevronRight", 13)}<span>${esc(name)}</span></nav>
|
| 791 |
<header class="tp-head ds-head">
|
| 792 |
-
<h1>${ownerLink(org, "/")}${esc(name)}</h1>
|
|
|
|
| 793 |
<p class="lede" id="sp-lede">${esc(describe(me))}</p>
|
| 794 |
<div class="facts sp-facts"><div class="sp-fi" id="sp-facts">${factsLine(me)}</div></div>
|
| 795 |
<div class="facts sp-links" id="sp-links">${linksLine(me)}</div>
|
|
@@ -1371,16 +1374,27 @@ async function openTask(me, i) {
|
|
| 1371 |
<button class="icon-btn sm" type="button" data-task-move="-1" aria-label="Previous task in this page" ${i === 0 ? "disabled" : ""}>${icon("chevronRight", 15, "flip")}</button>
|
| 1372 |
<button class="icon-btn sm" type="button" data-task-move="1" aria-label="Next task in this page" ${i + 1 === me.tasks.length ? "disabled" : ""}>${icon("chevronRight", 15)}</button>
|
| 1373 |
<button class="icon-btn sm" type="button" id="tk-close" aria-label="Close task">${icon("x", 15)}</button></div>
|
| 1374 |
-
<div class="tk-actions">${copyBtn(link.href, "Copy task link")}
|
| 1375 |
${rel ? `<a class="btn sm" href="/t/${enc(hub)}/${enc(rel)}" title="Open it in the explorer, where it runs on your account">${icon("database", 13)}Open in the explorer</a>` : ""}
|
| 1376 |
${canPlay ? `<button class="btn sm primary" type="button" data-play-task="${esc(idx)}">${icon("play", 13)}Start this task</button>` : ""}</div>
|
| 1377 |
<div class="tk-detail-b">${taskContent(t)}<p class="fine" id="tk-full-status" role="status">Loading the full taskβ¦</p></div>`;
|
| 1378 |
d.scrollIntoView({ block: "nearest", behavior: "smooth" });
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1379 |
try {
|
| 1380 |
const full = await api(`/api/spaces/${enc(me.spec)}/task?${new URLSearchParams({ env: me.taskPage.env, split: me.taskPage.split, index: idx })}`);
|
| 1381 |
if (page !== me || requestKey !== me.taskRequest || me.selectedTask !== i || !d.isConnected || d.hidden) return;
|
| 1382 |
if (full.task && typeof full.task === "object" && !Array.isArray(full.task)) {
|
| 1383 |
$(".tk-detail-b", d).innerHTML = taskContent(full.task);
|
|
|
|
| 1384 |
} else $("#tk-full-status", d).textContent = "Showing the task record returned by the list endpoint.";
|
| 1385 |
} catch (error) {
|
| 1386 |
if (page === me && requestKey === me.taskRequest && me.selectedTask === i && d.isConnected && !d.hidden)
|
|
|
|
| 782 |
if (!s) return;
|
| 783 |
const [org, name] = spec.split("/");
|
| 784 |
const sections = sectionsOf(me);
|
| 785 |
+
const linkedIndex = new URLSearchParams(location.search).get("task");
|
| 786 |
+
const taskView = /^\d{1,10}$/.test(linkedIndex || "");
|
| 787 |
const kind = me.L?.openenv_verified ? "OpenEnv Space" : s.declared_openenv || s.openenv || s.manifest ? "Unverified Space" : s.framework === "ors" ? "ORS Space" : "environment Space"; // as the server-rendered title says
|
| 788 |
+
if (!new URLSearchParams(location.search).has("task")) setMeta({ title: `${s.heading || name} Β· ${kind}`, description: `${spec}: an RL environment Space on Hugging Face. ${describe(me)} See it live: its app, a playground with rewards, its tasks, and MCP for coding agents.` });
|
| 789 |
me.viewer?.destroy();
|
| 790 |
me.viewer = null;
|
| 791 |
el.innerHTML = `<div class="wrap page${me.rendered ? "" : " fade-in"}">
|
| 792 |
<nav class="crumbs" aria-label="Breadcrumb"><a href="/">Environments</a>${icon("chevronRight", 13)}<a href="/?k=space">Spaces</a>${icon("chevronRight", 13)}<span>${esc(name)}</span></nav>
|
| 793 |
<header class="tp-head ds-head">
|
| 794 |
+
<h1>${taskView ? `Task ${Number(linkedIndex) + 1}` : `${ownerLink(org, "/")}${esc(name)}`}</h1>
|
| 795 |
+
${taskView ? `<p><a href="/s/${enc(spec)}">All tasks in ${esc(spec)}</a></p>` : ""}
|
| 796 |
<p class="lede" id="sp-lede">${esc(describe(me))}</p>
|
| 797 |
<div class="facts sp-facts"><div class="sp-fi" id="sp-facts">${factsLine(me)}</div></div>
|
| 798 |
<div class="facts sp-links" id="sp-links">${linksLine(me)}</div>
|
|
|
|
| 1374 |
<button class="icon-btn sm" type="button" data-task-move="-1" aria-label="Previous task in this page" ${i === 0 ? "disabled" : ""}>${icon("chevronRight", 15, "flip")}</button>
|
| 1375 |
<button class="icon-btn sm" type="button" data-task-move="1" aria-label="Next task in this page" ${i + 1 === me.tasks.length ? "disabled" : ""}>${icon("chevronRight", 15)}</button>
|
| 1376 |
<button class="icon-btn sm" type="button" id="tk-close" aria-label="Close task">${icon("x", 15)}</button></div>
|
| 1377 |
+
<div class="tk-actions"><a class="btn sm" href="${esc(link.href)}">Open task page</a>${copyBtn(link.href, "Copy task link")}
|
| 1378 |
${rel ? `<a class="btn sm" href="/t/${enc(hub)}/${enc(rel)}" title="Open it in the explorer, where it runs on your account">${icon("database", 13)}Open in the explorer</a>` : ""}
|
| 1379 |
${canPlay ? `<button class="btn sm primary" type="button" data-play-task="${esc(idx)}">${icon("play", 13)}Start this task</button>` : ""}</div>
|
| 1380 |
<div class="tk-detail-b">${taskContent(t)}<p class="fine" id="tk-full-status" role="status">Loading the full taskβ¦</p></div>`;
|
| 1381 |
d.scrollIntoView({ block: "nearest", behavior: "smooth" });
|
| 1382 |
+
const updateTaskMeta = (task) => {
|
| 1383 |
+
const q = new URLSearchParams(location.search);
|
| 1384 |
+
if (q.get("task") === String(idx) && q.get("env") === me.taskPage.env && q.get("split") === me.taskPage.split) {
|
| 1385 |
+
const prompt = [task.prompt, task.instruction, task.description, task.question].find((v) => typeof v === "string") || "";
|
| 1386 |
+
setMeta({ title: `${title} Β· ${me.spec}`, description: `${title}: ${me.taskPage.split} task in ${me.spec}, an RL environment on the Hugging Face Hub. ${prompt}` });
|
| 1387 |
+
const heading = $(".tp-head h1", me.el);
|
| 1388 |
+
if (heading) heading.textContent = title;
|
| 1389 |
+
}
|
| 1390 |
+
};
|
| 1391 |
+
updateTaskMeta(t);
|
| 1392 |
try {
|
| 1393 |
const full = await api(`/api/spaces/${enc(me.spec)}/task?${new URLSearchParams({ env: me.taskPage.env, split: me.taskPage.split, index: idx })}`);
|
| 1394 |
if (page !== me || requestKey !== me.taskRequest || me.selectedTask !== i || !d.isConnected || d.hidden) return;
|
| 1395 |
if (full.task && typeof full.task === "object" && !Array.isArray(full.task)) {
|
| 1396 |
$(".tk-detail-b", d).innerHTML = taskContent(full.task);
|
| 1397 |
+
updateTaskMeta(full.task);
|
| 1398 |
} else $("#tk-full-status", d).textContent = "Showing the task record returned by the list endpoint.";
|
| 1399 |
} catch (error) {
|
| 1400 |
if (page === me && requestKey === me.taskRequest && me.selectedTask === i && d.isConnected && !d.hidden)
|
web/js/util.js
CHANGED
|
@@ -1,14 +1,35 @@
|
|
| 1 |
// Shared helpers: DOM, formatting, API, markdown, modal, loading states.
|
| 2 |
import { icon, FILE_ICON } from "./icons.js";
|
|
|
|
| 3 |
|
| 4 |
export const $ = (s, el = document) => el.querySelector(s);
|
| 5 |
// the page's title, description and canonical address as you move around (the server writes them for a first load)
|
| 6 |
-
export function setMeta({ title, description } = {}) {
|
| 7 |
-
const
|
| 8 |
-
|
| 9 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
const canon = document.querySelector('link[rel="canonical"]');
|
| 11 |
-
if (canon) canon.href =
|
| 12 |
}
|
| 13 |
|
| 14 |
// go to another page of the app: a real path (/d/org/name, /t/...), so every page has its own address for search
|
|
|
|
| 1 |
// Shared helpers: DOM, formatting, API, markdown, modal, loading states.
|
| 2 |
import { icon, FILE_ICON } from "./icons.js";
|
| 3 |
+
import { SITE, DESCRIPTION, canonicalPath, imagePath } from "./seo-meta.js";
|
| 4 |
|
| 5 |
export const $ = (s, el = document) => el.querySelector(s);
|
| 6 |
// the page's title, description and canonical address as you move around (the server writes them for a first load)
|
| 7 |
+
export function setMeta({ title, description, indexable } = {}) {
|
| 8 |
+
const url = location.origin + canonicalPath(location.href);
|
| 9 |
+
const rendered = document.querySelector('meta[name="rlx-page-url"]');
|
| 10 |
+
const changed = rendered?.content !== url;
|
| 11 |
+
if (changed) {
|
| 12 |
+
document.querySelectorAll('script[type="application/ld+json"]').forEach((s) => s.remove());
|
| 13 |
+
if (rendered) rendered.content = url;
|
| 14 |
+
}
|
| 15 |
+
if (title || changed) document.title = !title ? `${SITE}: RL environments on the Hugging Face Hub` : title.includes(SITE) ? title : `${title} Β· ${SITE}`;
|
| 16 |
+
const desc = (description || (changed ? DESCRIPTION : document.querySelector('meta[name="description"]')?.content) || DESCRIPTION).slice(0, 300);
|
| 17 |
+
const set = (selector, value) => document.querySelector(selector)?.setAttribute("content", value);
|
| 18 |
+
set('meta[name="description"]', desc);
|
| 19 |
+
for (const key of ["og:title", "twitter:title"]) set(`meta[${key.startsWith("og:") ? "property" : "name"}="${key}"]`, document.title);
|
| 20 |
+
set('meta[property="og:description"]', desc);
|
| 21 |
+
set('meta[name="twitter:description"]', desc);
|
| 22 |
+
set('meta[property="og:url"]', url);
|
| 23 |
+
if (changed) {
|
| 24 |
+
const image = location.origin + imagePath(location.href);
|
| 25 |
+
set('meta[property="og:image"]', image);
|
| 26 |
+
set('meta[name="twitter:image"]', image);
|
| 27 |
+
set('meta[property="og:type"]', /^\/t\//.test(location.pathname) || location.search.includes("task=") ? "article" : "website");
|
| 28 |
+
}
|
| 29 |
+
if (changed || indexable !== undefined) set('meta[name="robots"]', indexable === false || /^\/(runs|run\/|compare\/)/.test(location.pathname)
|
| 30 |
+
? "noindex, follow" : "index, follow, max-image-preview:large, max-snippet:-1");
|
| 31 |
const canon = document.querySelector('link[rel="canonical"]');
|
| 32 |
+
if (canon) canon.href = url;
|
| 33 |
}
|
| 34 |
|
| 35 |
// go to another page of the app: a real path (/d/org/name, /t/...), so every page has its own address for search
|
web/rlx.css
CHANGED
|
@@ -395,6 +395,7 @@ body[data-view="home"] .top-search { visibility: hidden; }
|
|
| 395 |
html:not(.js) #view > .boot { display: none; }
|
| 396 |
.ssr ul { padding-left: 18px; display: grid; gap: 6px; }
|
| 397 |
.ssr .crumbs { margin-bottom: 12px; }
|
|
|
|
| 398 |
|
| 399 |
/* ββ RL framework rows (renderers/rl.js): NeMo Gym, Verifiers, verl / SkyRL βββββββββββββββββββββββββββββββββββββββ
|
| 400 |
A conversation as the policy gets it and the tools it may call. Nothing scrolls inside the page: long parts fold. */
|
|
|
|
| 395 |
html:not(.js) #view > .boot { display: none; }
|
| 396 |
.ssr ul { padding-left: 18px; display: grid; gap: 6px; }
|
| 397 |
.ssr .crumbs { margin-bottom: 12px; }
|
| 398 |
+
.seo-prompt { white-space: pre-wrap; overflow-wrap: anywhere; max-width: 85ch; line-height: 1.65; }
|
| 399 |
|
| 400 |
/* ββ RL framework rows (renderers/rl.js): NeMo Gym, Verifiers, verl / SkyRL βββββββββββββββββββββββββββββββββββββββ
|
| 401 |
A conversation as the policy gets it and the tools it may call. Nothing scrolls inside the page: long parts fold. */
|
web/social/rl-explorer.png
ADDED
|