LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Abstract
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
Community
Hi @akhaliq and the HF Papers team, thank you for featuring our paper!
Our teaser video is very wide (2.75:1), so the Daily Papers card crops out half of the result.
If possible, could you replace the media with the attached version, which is re-framed for the card?
If replacing it isn't possible, no worries at all. Please just keep the paper as it is.
Thank you!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation (2026)
- Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation (2026)
- Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting (2026)
- Exo2EgoHOI: Hand-Object-Interaction Aware Exocentric-to-Egocentric Video Generation (2026)
- SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models (2026)
- ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery (2026)
- 4Director: Controlling Video World Models with Rigid 3D Geometry (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper