From a deer’s motion trail to a spatial archive of the planet: a scientific reading of Bilawal Sidhu’s experiments, and the limits of 3D reconstruction.
Imagine that a video of a deer and its young does not disappear frame by frame. Instead, each moment leaves a three-dimensional trace: a turn of the head, the position of a body, the next step. You can move around these traces and read motion as a shape extending through space. The question changes from “What happened in this video?” to “Where did each moment happen, and how can we explore it?”
That is the starting point of Bilawal Sidhu’s video on X, posted on September 30, 2026, and also available on YouTube. Sidhu proposes a “spatial memory palace”: reconnecting photos and videos to the places and times in which they were captured, then assembling them into a scene we can navigate. Understanding this idea scientifically requires distinguishing the original observation, the geometry inferred by a model, and the way the result is displayed.

What does the video say?
Sidhu begins with the deer experiment and then applies the idea to footage of himself walking. He moves from seconds of motion to the same place across different seasons, photographs from a TED conference, a reconstruction combining an aircraft accident site with tracking data, and finally satellite imagery and storms. The common thread is turning separate files into a record connected through space and time.
| Chapter starts | Main idea |
|---|---|
| 00:45 | Defining 4D: three spatial dimensions plus time; leaving video moments inside a 3D scene. |
| 01:44 | A walking experiment and a literary analogy for imagining an entire human life as a temporal extension. |
| 02:32 | Reflections on free will and a cosmic record: philosophical and religious metaphors, not experimental findings. |
| 03:56 | Aligning winter footage with a spatial model captured in summer. |
| 04:43 | Connecting TED photographs to camera positions inside the venue and projecting them into space. |
| 06:09 | Combining a model of the Flight 7598 accident site with ground photographs and flight tracking. |
| 07:46 | A planetary archive: storm imagery, wind layers, and forecast tracks. |
| 09:53 | A shared archive in which people retain private memories and voluntarily contribute public material. |
The presentation explains a vision through visual experiments. It does not provide a published accuracy evaluation for every reconstruction, and it does not name the model used for the deer experiment. Watching the result alone cannot establish which algorithm produced all of it.
What does “four-dimensional” mean here?
A point in a three-dimensional scene has spatial coordinates: x, y, and z. Add the time at which it appears, t, and we have a spatiotemporal description: (x, y, z, t). We can select one time and inspect the scene at that instant, or display samples from several times together to reveal a movement path.

An animation analogy is onion skinning, which displays earlier or later poses alongside the current one so an artist can read movement. In this experiment, those poses become a spatial representation that can be viewed from different directions.
Visual representations of time have a longer history. Video Summagator, presented at CHI in 2012, offered a volumetric interface for exploring a video as a space-time cube. A video cube arranges two image dimensions and time; scene reconstruction also tries to infer spatial depth. Arranging images within a volume and estimating the structure of the photographed world are different operations.
Here, “4D” describes data and a computational model. It does not demonstrate beings that can see all of time, settle the question of free will, or enable physical travel into the past. Those ideas belong to the reflections prompted by the experiment and should be read separately from its technical result.
How does a flat image become a 3D scene?
An image records the color of each pixel, but does not automatically reveal the distance to the surface it represents. An object that appears small could be physically small or simply far from the camera. Reconstruction therefore needs additional geometric information or an estimate based on what a model has learned.
First, locate the camera. We need lens properties and the camera’s position and orientation relative to the scene. Traditional approaches match common details across overlapping photographs taken from different viewpoints. The COLMAP tutorial describes Structure-from-Motion recovering 3D structure and camera parameters, followed by Multi-View Stereo for a denser representation.
Second, estimate depth. Multi-view methods compare how details appear across images. Learned models can also predict geometry directly. VGGT, presented at CVPR 2025, infers camera parameters, depth maps, point maps, and 3D point tracks from a variable number of views. Depth Anything 3 estimates spatially consistent geometry from visual inputs with or without known camera poses. These are research examples that explain the field; the video does not state that either was used.
Third, establish a shared reference frame. With a pixel’s depth and camera parameters, it can be back-projected to a 3D position along the camera ray toward the surface. Points are then transformed from each camera’s coordinates into scene coordinates, producing a colored point cloud. Placing that reconstruction on a world map also requires scale, orientation, and geographic alignment.
Fourth, handle movement and time. A moving object changes position between images. Treating it as static can duplicate or distort its geometry. Motion needs time-associated points or poses, or a model describing how the scene changes. Combining temporal samples in a display is one option; by itself, this does not establish accurate geometric continuity between every instant.
Point clouds and Gaussian Splatting: what is the difference?
Many scenes in the video look like scattered points. Appearance alone cannot identify the reconstruction algorithm. A point cloud represents spatial samples and may include colors. 3D Gaussian Splatting represents a scene with Gaussian elements carrying size, orientation, appearance, and opacity, which are projected to render a requested view.
Kerbl and colleagues’ 2023 paper demonstrated high-quality novel views and fast rendering in its benchmarks. For changing scenes, 4D Gaussian Splatting, presented at CVPR 2024, combines 3D Gaussian elements with spatiotemporal information and a network predicting their deformation at different times.
| Concept | What it describes | What it does not guarantee |
|---|---|---|
| Point cloud | Spatial samples of a scene | Complete surfaces or details hidden from the cameras |
| Depth estimation | Estimated distance along each pixel’s viewing direction | Absolutely correct geometry for every detail |
| Gaussian Splatting | An appearance representation for rendering views | Survey-grade accuracy merely because the image is convincing |
| 4D representation | A 3D scene changing over time | A record of every event or knowledge of the future |
These ideas are related but serve different roles. A convincing rendering can still contain uncertain geometry.
When one place holds several times
One of the clearest examples connects winter footage to a spatial model captured in summer. Sidhu describes aligning the recording locations with the model and projecting video into the place. We can then move between memories of a location rather than browse rectangles in an album.

There is an important geometric challenge: places change. Trees grow, furniture moves, and snow or lighting can hide once-visible details. Aligning a site’s stable structure should therefore be distinguished from projecting elements recorded at another time. Combining them in one interface does not make their dates identical.
In the TED example, Sidhu describes using public photographs released under Creative Commons licenses alongside his own pictures, and connecting them to camera positions inside the venue. This makes it possible to search for a person across photographs, move backstage, or combine projections of speakers recorded at different moments.


Reprojecting an image using its depth can give it volume and permit a limited change of viewpoint. It does not automatically recover what was behind a person or outside the frame. As the viewing angle moves farther from the original camera angle, additional observations or inferred content become necessary. A useful spatial archive therefore retains the original photograph beside its reconstruction.
Reconstructing an event: context alone does not establish cause
In the Flight 7598 chapter, Sidhu says he used aerial footage released by the NTSB to reconstruct the accident site, then added flight tracking and ground photographs. The visual benefit is clear: the runway, motion traces, wreckage, and nearby pictures can be understood within one reference frame.

Overlaying sources does not synchronize them automatically. A ground photograph could have been captured hours after an accident, while tracking data may arrive at different time intervals. A line connecting two tracking samples may be interpolation for display, rather than a direct measurement of every intermediate position.
Three questions matter: Where did each element come from? When was it recorded? What was inferred or modeled? Organizing the available evidence can clarify context and help examine hypotheses, but a visual replay alone does not establish an accident’s cause.
A planetary memory: observations and forecasts are different layers
The video eventually reaches satellite images, storm tracks, and wind layers. This suits the idea of a broad spatiotemporal archive: select a region and date, then examine changes through successive measurements and images.

We must distinguish observations, derived from measurements and imagery, from forecasts, which are model outputs estimating a later state. NOAA explains that numerical weather prediction processes current observations within a model framework to produce future forecasts. Results can change when new measurements arrive.

A specific example of uncertainty is the U.S. National Hurricane Center’s operational forecast cone. It describes a range for the storm center’s track based on historical forecast errors, rather than the boundaries of every storm impact. Combining an observed past and a forecast future is useful when the meaning of each layer remains visible.
What makes a spatial archive scientifically trustworthy?
My proposed evaluation criteria go beyond attractive output. An inspectable archive needs original images or measurements, capture time distinct from upload time, a clear reference frame and scale, the model version, a distinction between observed and inferred elements, and a way to communicate error or limited confidence. When an estimate changes, the archive should retain an explanation of why.
The original God’s Eye View project documentation already distinguishes live data from simulated or approximate layers. For example, rendered traffic vehicles do not imply observation of every real vehicle. This illustrates a methodological point: real information can appear inside a scene that also contains modeled elements.
For personal memories, sharing must be a clear decision. A photograph can record a home, face, or private place; connecting photographs reveals more context than each image alone. Sidhu proposes retaining private material while optionally contributing public content. Turning that vision into a product requires understandable permissions and control over removal and sharing.
Why should designers and researchers care?
Potential applications include documenting places and heritage, comparing a site across distant dates, locating educational material in its setting, and building interactive memory experiences. For a multimedia designer, the work extends beyond arranging images: the interface must explain where the viewer stands, which time is being shown, the source of an element, and the uncertainty attached to it.
Sidhu’s experiment points toward a change in knowledge interfaces: from a file we search for in a folder to an event we reach through space and time. Its potential scientific value rests on sources that can be compared and inspected. Every place might eventually have an explorable archive; its trustworthiness begins with showing what we know, how we know it, and where the data ends.
About this article and its images: This is an original analytical article based on reviewing the video’s automated English captions, inspecting extracted frames, and comparing its claims with the references below. It is not a verbatim transcript. All eight images are original frames extracted from the video without generating or altering their content; visible in-frame credits are retained. Rights remain with their respective owners. The images illustrate the examples discussed here. A software license does not automatically cover video imagery or third-party data. This English edition accompanies the Arabic article on Ahmed Jalal’s website.
References
- Source video on X and YouTube.
- Nguyen, Niu & Liu — Video Summagator, CHI 2012.
- COLMAP — Official reconstruction tutorial.
- Wang et al. — VGGT, CVPR 2025.
- Lin et al. — Depth Anything 3, 2025.
- Kerbl et al. — 3D Gaussian Splatting, SIGGRAPH 2023.
- Wu et al. — 4D Gaussian Splatting, CVPR 2024.
- NOAA NCEI — Numerical weather prediction.
- NHC — Definition of the track forecast cone.
- Original God’s Eye View repository.
Sources reviewed: September 30, 2026.