Honeycomb: Constant-Size Scene Memory
Representation for Video World Models

Jack Wei Lun Shi2,*, Kaichen Zhou1,3,*,
Haoyu Chen1, Yufeng Weng2, Keane Ong2,3, Ruojin Cai1, Hang Hua4,
Justin K.W. Yeoh2, Mengyu Wang1
1 Harvard University 2 National University of Singapore
3 MIT 4 MIT-IBM Watson AI Lab

* Equal contribution.

Writing a video into HexMemory

Drag to rotate, scroll to zoom

HexMemory Visualization

Spatial Planes

front view XY
top-down XZ
side view YZ

Spatiotemporal Planes

ZT · walk → · τ ↑
YT · height → · τ ↑
XT · lateral → · τ ↑
|feature| (channel L2) never written extent before the last write points of the current chunk, written when the chunk ends

Abstract

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation.

Method

Overview of the Honeycomb pipeline
Overview of the Honeycomb pipeline. For scene-consistent video generation, we a) initialize a fixed-size HexMemory, b) read it to condition a generated chunk, and then c) write new observations back into the same tensors.
Memory design comparison: Spatia, LSM-World and Honeycomb
Memory design comparison. Spatia backprojects RGB observations into a point cloud, while LSM-World accumulates latent points. Honeycomb instead writes to HexMemory with fixed-size feature storage.
Point projection to a plane and bilinear splatting to four neighbouring cells
Bilinear splatting within the feed-forward writer. Left: point pi maps to location qi on plane P. Right: its learned contribution is distributed to the four neighbouring cells using bilinear weights wi,m. This operation is repeated for all six planes.

Closed-Loop Revisit

Revisiting a scene with constant-size memory: Honeycomb, Spatia and LSM-World
Revisiting a scene with constant-size feature memory. Given a single input frame and a camera trajectory that moves away and returns to the initial pose, Honeycomb generates a video whose final frame is consistent with the input frame. Videos below: Spatia is shown in its native 24 FPS format, while LSM-World and Honeycomb are shown at native 16 FPS.
HoneycombSpatiaLSM-World

Novel View Synthesis

Honeycomb generates 3 autoregressive chunks from a single input image along the RealEstate10K camera trajectory. Select an example to play it.

Quantitative Results

Evaluation results on WorldScore. The Average Score is the mean of the Static and Dynamic Scores; all remaining metrics are computed by the WorldScore benchmark.
MethodAverage
Score
Static
Score
Dynamic
Score
3D
Const
Photo
Const
Style
Const
Subject
Quality
Models with 3D cache
WonderJourney54.1963.7544.6380.6079.0362.8266.56
WonderWorld61.7972.6950.8886.8785.5670.5749.81
Spatia63.2164.8861.5483.2689.0983.3346.66
LSM-World61.2062.6959.7080.8876.10––
General video models
VideoCrafter250.0352.5747.4965.1461.8543.7956.74
EasyAnimate52.2552.8551.6567.2947.3573.0550.31
Allegro53.6455.3151.9770.5069.8965.6047.41
Wan2.155.2157.5652.8578.7478.3677.1859.38
Honeycomb65.5268.0163.0382.2985.7684.2146.28
Novel-view synthesis on RealEstate10K and closed-loop on WorldScore. We report evaluation results of all baseline methods using their default settings.
RE10K NVSWorldScore closed-loop
MethodPSNR↑SSIM↑LPIPS↓ PSNRC↑SSIMC↑LPIPSC↓FlowC↓
ViewCrafter12.280.5120.57112.320.3690.57430.78
FlexWorld13.170.5670.54412.860.4300.60255.77
Voyager14.670.5770.49315.990.4590.4237.11
Spatia15.580.6160.39015.670.4880.3536.64
LSM-World17.460.6360.45215.120.4600.46327.05
Honeycomb18.450.6740.27417.220.5040.3113.00

Conclusion

Maintaining scene consistency over long video rollouts requires persistent memory that can efficiently incorporate new observations. In this work, we introduce Honeycomb, a video world model with HexMemory. This design keeps feature storage fixed without reprocessing the entire history at each write. Experiments on WorldScore and RealEstate10K demonstrate better generation quality, novel-view synthesis, and revisit consistency. These results highlight the potential of recurrent feature memory for efficient and consistent video world modeling.