AI Stack Exchange
2026-09-12 12:40 UTC
By Sattyam Jain
AI-110-20260912-social-media-f5b6f65c
What is the correct denominator for Attack Success Rate when the attack is search-based?
Attack Success Rate is usually reported as (successful attacks) / (attempted attacks). For a fixed attack set that is unambiguous. For a search-based attack, it is not, and I cannot find a treatment of this. If my attack is a search over instruction perturbations with a budget of N queries per task, then: Denominator = number of tasks: ASR goes up monotonically with N, because more search finds more. Two papers with different budgets are not comparable, and neither paper has done anything wrong. Denominator = number of queries: ASR goes down monotonically with N, because most queries in a large budget fail. Also not comparable. Denominator = number of tasks, with N reported as a parameter: comparable only between papers that happen to pick the same N. Questions: Is there an accepted convention in the adversarial-ML literature for reporting a budget-dependent success rate? Something like "success at budget N" reported as a curve rather than a scalar? Is there an existing name for the scalar summary of such a curve? Area under it, or the budget at which success first exceeds a threshold, feel like the obvious candidates, but I would rather use an existing term. For hypothesis testing on such a rate, does the search budget need to enter the interval, or is a Wilson interval on (successes/tasks) at a fixed N defensible? Context: I work on adversarial evaluation of robot policies, where the same question arises, and the literature reports bare scalars.
Attack Success Rate is usually reported as (successful attacks) / (attempted attacks). For a fixed attack set that is unambiguous. For a search-based attack, it is not, and I cannot find a treatment of this. If my attack is a search over instruction perturbations with a budget of N queries per task, then: Denominator = number of tasks: ASR goes up monotonically with N, because more search finds more. Two papers with different budgets are not comparable, and neither paper has done anything wrong. Denominator = number of queries: ASR goes down monotonically with N, because most queries in a large budget fail. Also not comparable. Denominator = number of tasks, with N reported as a parameter: comparable only between papers that happen to pick the same N. Questions: Is there an accepted convention in the adversarial-ML literature for reporting a budget-dependent success rate? Something like "success at budget N" reported as a curve rather than a scalar? Is there an existing name for the scalar summary of such a curve? Area under it, or the budget at which success first exceeds a threshold, feel like the obvious candidates, but I would rather use an existing term. For hypothesis testing on such a rate, does the search budget need to enter the interval, or is a Wilson interval on (successes/tasks) at a fixed N defensible? Context: I work on adversarial evaluation of robot policies, where the same question arises, and the literature reports bare scalars.
Full article content could not be extracted automatically. Read the original below.
Source:
AI Stack Exchange
· ai.stackexchange.com