- likes
- 40
- comments
- 1
Post
Imagine a robot arm moving a cube. As the arm moves, the cube disappears into thin air. It’s still there in reality, but the model no longer sees it. What happened? Most visual models see a scene as a collection of patches, rather than objects. They can capture multiple objects or just random parts of the same object. So the model is only as good as its patches. But what if we perceive the world as objects to begin with? 3D-DLP (published by CMU and Lambda at ICML 2026) learns the world as objects. Each object carries its position, size, and appearance, learned without the need for segmentation or labeling. Result: a robot moving a cube tracks that cube the whole way. It doesn't vanish because the tokenizer dropped it. Bonus: the model becomes more parameter efficient, has better generalization, and can generate semantically rich object boundaries. Read the paper: https://lnkd.in/gr5QuuYp