logo

NJP

‘No, Of Course I Can!’: Deeper Fine-Tuning Attacks That Bypass Token-Level Safety

ServiceNow · Apr 18, 2026 · video

Fine-tuning APIs offered by providers like OpenAI and Anthropic allow customers to customize frontier language models but they also introduce a subtle and underappreciated attack surface. Prior work has shown that safety alignment is “shallow,” concentrated in the first few response tokens. In this talk, Abhay Puri shows that existing fine-tuning attacks are equally shallow, and can therefore be blocked by simple token-level defenses that enforce an aligned refusal prefix. But what happens when attackers go deeper? We introduce NOICE (No, Of Course I Can Execute), a novel fine-tuning attack that flips the script: instead of training models to skip the refusal, it trains them to refuse first and then comply anyway. This “refuse-then-comply” strategy exploits the very mechanism safety systems rely on, deceiving output filters like LlamaGuard that treat an initial refusal as a signal of safety. The attack requires no overtly harmful training data, costs under $100 in API credits, and achieves attack success rates of 57% against GPT-4o and 72% against Claude Haiku seven times higher than prior harmless-data attacks. It earned a $2,000 bug bounty from OpenAI and was acknowledged as a novel vulnerability by Anthropic.

View original source

https://www.youtube.com/watch?v=abe9FXr7eEc