I'm unconvinced that the fact they're getting better at following instructions, like putting objects where the prompter specifies, or changing the colour, or putting the right number of them, etc means the model actually understands what the objects mean beyond their appearance. It doesn't understand the cultural meanings attached to each object, and thus is unable to truly make a decision about why it should place an apple rather than an orange, or how the message within the picture changes when it's a red sports car rather than a beige people-carrier.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: