There are multiple major flaws with watermarks for texts:
Their method only works for longer texts as it relies on statistical effects that aren't clear enough to detect in a very short text.
They are resting their approach on an incorrect assumption: their models prefer certain words and if they appear more frequently than in a random text, they assume it's generate by their model. But who says texts are random? Out of billions of people there will be some who have a similar preference for some words and their texts will always be wrongfully accused of AI-generated just because they happen to have a similar word choice preference. The shorter the texts the likelier that issue becomes.
There will soon be tools that replace a random amount of words with synonyms automatically and therefore remove any chances of detecting the watermark.