I remember seeing an article where workers had to write skill.md files for their specific domain knowledge. There was an AI tool that redacted the most specialized domain knowledge before turning in the skill.
Generally if the AI is an API and not on-prem you can find ways to leak huge amounts of data by sending it to the AI, etc.
Have the AI review its own work without pointing out particular problems you think are likely given your experience. In general give it as little context as possible, give vague instructions that could plausibly mean multiple things. Simple sabotage manual basically applies.