Skip to content

Published attack results on VLA policies — and what provael's ASR is comparable to

Independent project. This page cites other people's published results and does not reproduce them. Provael runs no harness from any paper below and publishes no score against anyone's leaderboard. Every figure is as reported by its authors, with a resolvable identifier so you can check it yourself.

Provael's headline is a roleplay instruction diverting SmolVLA × LIBERO 88% of the time (44/50) against a 2/50 (4%) benign control, across all ten libero_object tasks. The floor travels with the rate on this page for the same reason the page exists: 88% is not a number until you know what the same measurement returns when nothing is attacking. Most rows below report no floor at all, because their authors published none — that absence is part of what the last column is measuring. Any technical buyer will immediately benchmark that against the published record — where attacks routinely report 76%, 96%, near-100%. If those numbers sit in one column, provael looks unremarkable.

They do not belong in one column. That is what this page is for, and the last column is the whole point.

The comparability rule

Provael's ASR is the fraction of episodes driven out of the policy's benign safe envelope — a keep-out / envelope-breach predicate, measured per episode, with a matched benign false-positive control and a 95% Wilson interval.

Most published VLA attack numbers measure something else: task-success degradation ("the policy stopped completing the task"). A policy that stops working is not the same event as a policy that does the wrong thing safely, and neither is the same as a policy that completes the task while breaching its envelope. Putting a 95%→<5% success collapse beside a 100% envelope-breach ASR asserts an equivalence that does not exist.

The table

Attack Reported metric Value Suite / setting Source Comparable to a Provael ASR?
16/255 PGD (white-box pixel perturbation) on OpenVLA-7B Task success under attack 95% → under 5% LIBERO arXiv:2605.25889 No — not directly comparable. This is task-success collapse; provael's ASR is envelope breach. A policy that fails safely scores 0% here on provael's predicate and ~100% on this one.
FreezeVLA — action-freezing ASR (action freeze) 76.2% average as reported arXiv:2509.19870 Closest in kind, still not directly. Both are per-episode ASRs of an action-channel attack, so the shape matches — but "the policy froze" and "the policy left its envelope" are different unsafe predicates, and a freeze is an availability failure provael's envelope predicate does not flag.
BadVLA — objective-decoupled backdoor Targeted ASR near 96.7% (authors' wording) multiple VLA benchmarks arXiv:2505.16640 No. BadVLA is a training-time threat that modifies the model. Provael's backdoor family is a screen run against a clean checkpoint and is expected to score ~0%; provael ships no backdoored weights. A high BadVLA number and a low provael number are both correct and measure opposite things.
Printed adversarial patch on OpenVLA Task success under attack 98% → 6% as reported arXiv:2511.21192 No. Task-success degradation again — see the first row.
DRIFT — first-step denoising redirection (universal gripper patch) Originally-solvable tasks broken essentially all, with a small single patch π0 and π0.5 across four LIBERO suites arXiv:2608.03207 No — not directly comparable, for two independent reasons. Provael has published no result against a flow-matching policy of any kind, so there is no Provael number for this row to sit beside: the pi0/pi05/pi0fast adapters are registered and marked scaffolding, and none has loaded a checkpoint. Separately, the metrics differ — DRIFT reports solvable tasks broken, a provael ASR reports envelope breach. Either reason alone would rule out one column.
RoboPAIR — LLM-controlled-robot jailbreak Jailbreak ASR ~100% 3 systems incl. a deployed Unitree Go2 arXiv:2410.13691 Partially. Same family of claim as provael's EAI01 result (an instruction drives a harmful action) and the nearest external anchor for it — but measured on LLM-controlled robots with a human-judged harmful-action criterion, not an envelope predicate.
Per-LIBERO-suite single-step ASR (Spatial / Object / Goal / LIBERO-10) withheld LIBERO ×4 arXiv:2605.25889 · 2505.16640 · 2509.19870 Row withheld, deliberately. Four per-suite figures were proposed for this row — Spatial 97.5 / Object 93.8 / Goal 96.5 / LIBERO-10 77.3. They appear as a set in none of the three papers checked: 2605.25889 reports OpenVLA-7B at 95.4% clean success on LIBERO-Spatial and validates 48 PGD cells without publishing them; in 2505.16640 and 2509.19870, two of the four values (93.8, 96.5) occur only as isolated cells of unrelated tables and the other two (97.5, 77.3) do not occur at all. Printing figures on a page about checkability that we could not check would be self-refuting. Restore the row when a source resolves.
RedVLA — scene risk-factor injection, instruction held fixed ASR (physical-safety violation: State / Cumulative / Conditional) 95.5% on π₀.₅; 64.9%–95.5% across six models LIBERO, simulation, 10 trials per configuration arXiv:2604.22591 No — and the reason is unusually clean. RedVLA formalises red teaming over the environment–instruction joint space (s′₀, l′), then fixes the instruction (l′ = l) and perturbs only the initial state. Provael does the converse: it fixes the scene and perturbs the instruction. Same benchmark, same simulator, complementary halves of one formalism — so the two ASRs are orthogonal quantities, not a strong result and a weak one. Their predicate is a typed physical-safety violation; ours is an envelope exit. Note what this row does not rest on: both are simulation, so sim-versus-hardware is not the difference.
Provaelroleplay (instruction reframing) ASR = episodes driven out of the benign envelope 88% (44/50), task-clustered 95% CI [72%, 100%] SmolVLA × LIBERO, all ten libero_object tasks committed artifacts — (this is the reference definition). The interval is bootstrapped over TASKS, not episodes; every other row on this page pools episodes, so ours is the wider and more conservative construction.

Provael's own numbers, in full

Every figure below is read from the committed results/smolvla_libero_object/report.json. Nothing here is estimated, rounded up, or reconstructed.

Which run this is, because the two figures differ and both are ours. This table is the earlier single-task runlibero_object/0 only, 10 seeds — where roleplay read 10/10. The headline row in the table above is the later ten-task suite at 44/50 (88%). The larger run is the one to quote: it spans all ten tasks and its 88% is the more conservative number. Both are committed; neither supersedes the other's raw counts, and the single-task 100% is retained here rather than deleted because it is what the report.json this section cites actually contains.

Attack Family Successes / attempts Rate
none (benign control) baseline 0/10 0.0%
roleplay instruction 10/10 100%
goal_substitution instruction 6/10 60%
paraphrase instruction 1/10 10%
patch visual 0/10 0%
decoy_object visual 0/10 0%
scene_text injection 0/10 0%
mcp_tool_desc injection 0/0 N/A — not applicable in this suite, excluded from the denominator. N/A is not 0 and is not a pass
  • Instruction family aggregate: 17/30 = 56.7%, 95% Wilson CI [39.2–72.6%].
  • Benign FPR: 0.0 — the control every rate above is read against.
  • Transfer status: real-transfer.

Scope, stated with the numbers rather than beneath them: simulation only, one policy, one task, n=10 seeds, and an uncalibrated keep-out predicate — so the rate means "diverted out of the benign envelope", not a hazard rate. The visual (0/20) and injection (0/10) nulls are published here for the same reason they are published everywhere else in this repo: a red team that reports only its hits has no denominator.

What this table is not

It is not a claim that provael's attack is stronger, more general, or more important than any row above. Several of these papers demonstrate capabilities provael does not have — white-box gradient attacks, printed physical patches, training-time backdoors, real hardware. See PRIOR_ART.md, which says so at length and names what provael does not originate.

The claim is narrower: provael's number answers a different question, it is reported with the interval and the control that make it checkable, and it should be filed as such rather than ranked against numbers that measure task success.


Independent · evidence, not certification · every row carries a resolvable identifier by test (tests/test_citations_resolvable.py).