# Hugging Face Space Deployment This repository is prepared for a Gradio-based Hugging Face Space with ZeroGPU. ## Runtime - Space SDK: Gradio - Space hardware: ZeroGPU - Space Python: 3.10.13 - Public port: `7860` - Entrypoint: `python app.py` - Recommended use: ZeroGPU for request-scoped GPU allocation - In the Space settings UI, select `ZeroGPU` as the hardware target ## Model Assets The app first checks local model files under `LANCE_MODEL_BASE_DIR`. Default behavior: - Local checkout with `downloads/`: use `./downloads` - Hugging Face Space without local assets: download from `bytedance-research/Lance` into `/data/lance_models` - Video tasks use the pre-fetched `Lance_3B_Video` assets when available. - Startup prefetch downloads the model snapshots on CPU so the first GPU request does not pay that cold-start cost. - RIFE interpolation is optional. The app now falls back to the original video if the RIFE script or checkpoint is missing or incompatible. To restore interpolation, keep the RIFE code and `RIFE/train_log/flownet.pkl` from the same release. - Image tasks unload the active video model first, then load `Lance_3B`. - Switching back to a video task unloads `Lance_3B`, then reloads `Lance_3B_Video`. Useful environment variables: - `LANCE_MODEL_REPO_ID`: Hugging Face model repo to download from. Default: `bytedance-research/Lance` - `LANCE_MODEL_BASE_DIR`: directory containing `Lance_3B_Video`, `Qwen2.5-VL-ViT`, and `Wan2.2_VAE.pth` - `LANCE_VIDEO_MODEL_PATH`: explicit video model directory override - `LANCE_IMAGE_MODEL_PATH`: explicit image model directory override - `LANCE_MODEL_PATH`: legacy explicit model directory override used for both task families if the family-specific override is unset - `LANCE_MODEL_VARIANT`: `video` or `image`; default is `video` - `LANCE_AUTO_DOWNLOAD`: set to `1` to download missing assets from the Hub - `LANCE_GPUS`: comma-separated GPU IDs, for example `0` or `0,1` - `LANCE_QUEUE_SIZE`: Gradio queue size - `LANCE_GRADIO_TMP_ROOT`: output and temporary file directory - `LANCE_ZEROGPU_MAX_DURATION_SECONDS`: upper bound for the task-aware `@spaces.GPU` duration request in seconds (default cap: 300) - `LANCE_INSTALL_FLASH_ATTN_ON_STARTUP`: set to `1` to install the pinned flash-attn wheel during Space startup instead of inside the GPU reservation (the wheel matches Python 3.10.13 and torch 2.8.0) - `LANCE_PREFETCH_MODEL_ASSETS`: set to `0` to skip CPU-side model prefetch at startup - `LANCE_PREFETCH_MODEL_VARIANTS`: comma-separated model variants to prefetch, for example `video,image` Expected model layout: ```text ${LANCE_MODEL_BASE_DIR}/ Lance_3B_Video/ llm_config.json model.safetensors tokenizer.json ... Lance_3B/ llm_config.json model.safetensors tokenizer.json ... Qwen2.5-VL-ViT/ config.json vit.safetensors Wan2.2_VAE.pth ``` ## Local Docker Check ```bash docker build -t lance-space . docker run --gpus all -p 7860:7860 \ -e LANCE_MODEL_BASE_DIR=/models/lance \ -v /path/to/lance/downloads:/models/lance \ lance-space ``` Open `http://localhost:7860`. ## Files Not Uploaded The Space build excludes generated or heavyweight local files through `.dockerignore`: - `downloads/` - `results/` - `tmps/` - Python cache files