Study: LLMs Respond Differently To Harmful Prompts When AI Watermarking Is Present
New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. – Ars Technica
Read more »