Multimodal
Models that accept more than text — images, audio, video — in the same conversation. Image inputs are tokenized too, and cost real money per image.
Models that accept more than text — images, audio, video — in the same conversation. Image inputs are tokenized too, and cost real money per image.