Jeff Hawkins' Thousand Brains Theory proposes that the neocortex consists of many parallel predictive models, each operating at different timescales and levels of abstraction. Predictions flow down the cortical hierarchy; prediction errors flow up. Learning is driven entirely by the discrepancy between what was predicted and what was observed.
This architecture implements that principle for video understanding. A frozen visual backbone encodes frames; a three-level learned hierarchy compresses time at multiple scales, predicts future states in latent space (never pixels), and uses top-down context from abstract levels to constrain concrete predictions.
The approach shares principles with LeCun's JEPA (Joint Embedding Predictive Architecture), particularly the commitment to latent-space prediction over pixel reconstruction. Where V-JEPA 2 operates at a single temporal scale, this project builds the multi-scale temporal hierarchy that both Hawkins and LeCun theorized but neither has fully implemented.