Discussion about this post

User's avatar
Kai Williams's avatar

I’m working on a similar explainer (a year later haha) for a slightly more general audience, so I really appreciate this!

What’s your basis for assuming that trajectories are less dense token-wise than text is? Before reading your take here, I’d assumed the “pictures are worth a thousand words” dictum that video data would be much richer, all else being equal. Is the challenge that much of this data is extraneous (e.g. the textures of a flower on the window sill don’t matter if the robot is training to wash dishes)? Or that diversity is much more difficult to come by in text? Has someone tried to measure this?

Ludek Cizinsky's avatar

Wow, super interesting and resourceful article!

6 more comments...

No posts

Ready for more?