Lance / SPACE_DEPLOYMENT.md
ffy2000's picture
Vendor RIFE into repo
35616fd
|
Raw History Blame Contribute Delete
3.29 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade

Hugging Face Space Deployment

This repository is prepared for a Gradio-based Hugging Face Space with ZeroGPU.

Runtime

  • Space SDK: Gradio
  • Space hardware: ZeroGPU
  • Space Python: 3.10.13
  • Public port: 7860
  • Entrypoint: python app.py
  • Recommended use: ZeroGPU for request-scoped GPU allocation
  • In the Space settings UI, select ZeroGPU as the hardware target

Model Assets

The app first checks local model files under LANCE_MODEL_BASE_DIR.

Default behavior:

  • Local checkout with downloads/: use ./downloads
  • Hugging Face Space without local assets: download from bytedance-research/Lance into /data/lance_models
  • Video tasks use the pre-fetched Lance_3B_Video assets when available.
  • Startup prefetch downloads the model snapshots on CPU so the first GPU request does not pay that cold-start cost.
  • RIFE interpolation is optional. The app now falls back to the original video if the RIFE script or checkpoint is missing or incompatible. To restore interpolation, keep the RIFE code and RIFE/train_log/flownet.pkl from the same release.
  • Image tasks unload the active video model first, then load Lance_3B.
  • Switching back to a video task unloads Lance_3B, then reloads Lance_3B_Video.

Useful environment variables:

  • LANCE_MODEL_REPO_ID: Hugging Face model repo to download from. Default: bytedance-research/Lance
  • LANCE_MODEL_BASE_DIR: directory containing Lance_3B_Video, Qwen2.5-VL-ViT, and Wan2.2_VAE.pth
  • LANCE_VIDEO_MODEL_PATH: explicit video model directory override
  • LANCE_IMAGE_MODEL_PATH: explicit image model directory override
  • LANCE_MODEL_PATH: legacy explicit model directory override used for both task families if the family-specific override is unset
  • LANCE_MODEL_VARIANT: video or image; default is video
  • LANCE_AUTO_DOWNLOAD: set to 1 to download missing assets from the Hub
  • LANCE_GPUS: comma-separated GPU IDs, for example 0 or 0,1
  • LANCE_QUEUE_SIZE: Gradio queue size
  • LANCE_GRADIO_TMP_ROOT: output and temporary file directory
  • LANCE_ZEROGPU_MAX_DURATION_SECONDS: upper bound for the task-aware @spaces.GPU duration request in seconds (default cap: 300)
  • LANCE_INSTALL_FLASH_ATTN_ON_STARTUP: set to 1 to install the pinned flash-attn wheel during Space startup instead of inside the GPU reservation (the wheel matches Python 3.10.13 and torch 2.8.0)
  • LANCE_PREFETCH_MODEL_ASSETS: set to 0 to skip CPU-side model prefetch at startup
  • LANCE_PREFETCH_MODEL_VARIANTS: comma-separated model variants to prefetch, for example video,image

Expected model layout:

${LANCE_MODEL_BASE_DIR}/
  Lance_3B_Video/
    llm_config.json
    model.safetensors
    tokenizer.json
    ...
  Lance_3B/
    llm_config.json
    model.safetensors
    tokenizer.json
    ...
  Qwen2.5-VL-ViT/
    config.json
    vit.safetensors
  Wan2.2_VAE.pth

Local Docker Check

docker build -t lance-space .
docker run --gpus all -p 7860:7860 \
  -e LANCE_MODEL_BASE_DIR=/models/lance \
  -v /path/to/lance/downloads:/models/lance \
  lance-space

Open http://localhost:7860.

Files Not Uploaded

The Space build excludes generated or heavyweight local files through .dockerignore:

  • downloads/
  • results/
  • tmps/
  • Python cache files