Vision-Language-Action model — a single network that maps camera images and a language instruction directly to robot motor commands.
Continue to AI University →