God I hate these fake science calculation posts.
His prompts were each using 2.9 million token. That's a massive amount, he was basically purposefully using complex tasks that needed massive amounts of data parsing. We're talking about like 1 percent of users that are using 2.9 million per prompt.
The method used to calculate doesn't include batching. These companies aren't running one request per gpu here. Even the authors of the method admit it over estimates by 4 to 20x.
all 13 comments