apparently audio and images are more efficient compared to text for multimodal models?
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies:
apparently audio and images are more efficient compared to text for multimodal models?