ai security

PGD: Same Attack, Way More Persistent

Wed Aug 19 2026

This post includes proof-of-concept material for educational purposes. Only run it against systems you own or have explicit permission to test.

Last time, I threw the cheapest attack possible at a ResNet18. One gradient step, sign function, done. It flipped the prediction fine, but aimed at a specific target class, it whiffed completely. Never landed, not even at a perturbation big enough to see. This time, same model, same target, same dog. It lands by step five, locked at 100% confidence.

Here’s what changed. FGSM computes the gradient sign once and stops. This attack computes it, takes a small step, and does it again, forty times, correcting course as the loss landscape curves beneath it. Same photo, same target class, “toaster.” Run FGSM’s targeted variant at that same target and it never gets there, across its whole epsilon range. This one reaches it at the smallest budget I tested, five steps into forty available. This isn’t a stronger bug or a different exploit. It’s the same idea Goodfellow’s team described in 2014, just not stopping after one step. There’s a name for it too: PGD, Projected Gradient Descent, from a 2017 paper by Madry, Makelov, Schmidt, Tsipras, and Vladu. The core claim: don’t trust a linear approximation of the loss all the way to the edge of your budget in one jump. Trust it locally, take a small step, recompute the gradient, and project back onto the budget every time a step might have carried you past it.

If you’ve done real fuzzing or iterative red-teaming, this won’t be a surprising result. Looping with feedback beats a single blind shot, and that’s the entire reason fuzzers loop instead of guessing once. Same constraint as FGSM: white-box, gradient access required. This attack just costs more to run, forty forward and backward passes instead of one, so it’s not what you reach for when you need a five-second smoke test. It’s what you reach for when you need the real answer.

The idea, without the paper’s notation

FGSM’s whole attack is one step: compute the gradient sign, scale by epsilon, done. That’s optimal if you trust the linear approximation of the loss all the way out to the edge of your epsilon budget, but for anything beyond a tiny budget, that trust is misplaced. Loss landscapes curve. A straight line to the edge of the L∞ ball is rarely the direction that actually maximizes (or minimizes, for targeted attacks) loss once you’re partway there.

PGD’s fix: don’t try to get there in one jump. Take a small step of size alpha (a fraction of epsilon, epsilon / 10 by default here), recompute the gradient at the new point, take another step. Each individual step trusts the linear approximation only locally, where it’s actually valid. After every step, project back onto the epsilon-ball around the original image:

x_adv = clip(x_adv + step, x_orig - epsilon, x_orig + epsilon)

That projection is the whole reason this is called projected gradient descent. Without it, forty accumulated small steps would wander arbitrarily far from the original image. With it, the attack can explore, curve, correct course, and still never exceed the budget you set.

What this looks like in code

The loop, trimmed to the actual attack logic (full file linked below):

for step in range(n_steps):
    x_adv.requires_grad_(True)

    x_norm = normalize(x_adv)
    logits = model(x_norm.unsqueeze(0))
    loss = F.cross_entropy(logits, torch.tensor([label]).to(DEVICE))

    model.zero_grad()
    loss.backward()

    with torch.no_grad():
        if targeted:
            step_direction = -alpha * x_adv.grad.sign()
        else:
            step_direction = alpha * x_adv.grad.sign()

        x_adv = x_adv + step_direction

        # projection: clip back onto the epsilon-ball around x_orig
        x_adv = torch.max(torch.min(x_adv, x_orig + epsilon), x_orig - epsilon)
        x_adv = torch.clamp(x_adv, 0, 1)

Two details worth calling out because they’re the parts people get wrong on a first implementation:

Running it

git clone https://github.com/BogiLoco/PGD
cd PGD
python3 -m venv venv && source venv/bin/activate
pip install torch torchvision matplotlib pillow requests
python pgd.py

No training, no dataset download: same pretrained ResNet18 as FGSM, same sample image on first run. It runs an untargeted attack, a targeted attack, then an epsilon sweep for both.

An untargeted attack at eps=8/255 takes the same Samoyed at 83.72% confidence straight to lynx at 100.00% confidence. Not a dip and recovery like FGSM’s version of this attack. It walks straight there and locks in.

Then the targeted attack, the one FGSM couldn’t do. Same target, "toaster", eps=16/255, watching the per-step log:

PGD targeted -> toaster (eps=16/255):
  step   0: hen                       (5.2%)
  step   5: toaster                   (100.0%) <- target reached!
  step  10: toaster                   (100.0%) <- target reached!
  step  15: toaster                   (100.0%) <- target reached!
  step  20: toaster                   (99.7%) <- target reached!
  step  25: toaster                   (99.9%) <- target reached!
  step  30: toaster                   (100.0%) <- target reached!
  step  35: toaster                   (100.0%) <- target reached!
  step  39: toaster                   (100.0%) <- target reached!

Five steps to lock in. Here’s what that looks like:

Side-by-side: a Samoyed classified correctly at 83.72%, the PGD perturbation visually amplified, and the same image after a targeted PGD attack, classified as toaster at 100.00% confidence

Same caveat as FGSM: this runs against a resnet18 instance in your own process, on an image you provide. Still a white-box attack, gradient access required, so it says nothing about a model you can only reach through an API.

Full epsilon sweeps for the same image, untargeted and targeted:

   epsilon | prediction                | confidence
----------------------------------------------------
    0.0000 | Samoyed                   |    83.72%
    0.0039 | white wolf                |    98.75%  <- attack successful!
    0.0078 | white wolf                |    96.52%  <- attack successful!
    0.0157 | timber wolf               |    99.94%  <- attack successful!
    0.0314 | lynx                      |   100.00%  <- attack successful!
    0.0627 | lynx                      |   100.00%  <- attack successful!
    0.1255 | Siamese cat               |   100.00%  <- attack successful!
   epsilon | prediction                | confidence
----------------------------------------------------
    0.0000 | Samoyed                   |    83.72%
    0.0039 | toaster                   |    84.53%  <- SUCCESS: model now sees the target!
    0.0078 | toaster                   |   100.00%  <- SUCCESS: model now sees the target!
    0.0157 | toaster                   |   100.00%  <- SUCCESS: model now sees the target!
    0.0314 | toaster                   |   100.00%  <- SUCCESS: model now sees the target!
    0.0627 | toaster                   |   100.00%  <- SUCCESS: model now sees the target!
    0.1255 | toaster                   |   100.00%  <- SUCCESS: model now sees the target!

The last post didn’t walk through FGSM’s targeted sweep in detail, but it’s in the repo and worth running yourself: FGSM never reaches "toaster" at all, not even at eps=64/255. This one reaches it at eps=1/255, the smallest nonzero budget tested, already at 84.53% confidence there. By eps=2/255 it’s saturated at 100%.

Before you go

Sources