[refuse] vs [respond]: The two-token trick that simplifies LLM safety
A simple token-level method for controlling when and why language models refuse to answer, without retraining.
A simple token-level method for controlling when and why language models refuse to answer, without retraining.