Genie: Learning Interactive World Models from Unlabeled Images
Labeling 30,000 hours of video is an expensive path, and the bill scales fast. This article breaks down how the Latent Action Model (LAM) learns interaction from unlabeled images, and what that means economically and technically for agent training.