Model Behavior — AI news1 storiesWatermarked LLMs Sometimes Follow Harmful Instructions Ars Technica AITopicsArchive