LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse.
This is a summary curated by AIFuture. Read the complete article at the original source:
Read the full story on Ars Technica