Attack Success Rate is usually reported as (successful attacks) / (attempted attacks). For a fixed attack set that is unambiguous. For a search-based attack, it is not, and I cannot find a treatment of this. If my attack is a search over instruction perturbations with a budget of N queries per task, then: Denominator = number of tasks: ASR goes up monotonically with N, because more search finds more. Two papers with different budgets are not comparable, and neither paper has done anything wrong. Denominator = number of queries: ASR goes down monotonically with N, because most queries in a large budget fail. Also not comparable. Denominator = number of tasks, with N reported as a parameter: comparable only between papers that happen to pick the same N. Questions: Is there an accepted convention in the adversarial-ML literature for reporting a budget-dependent success rate? Something like "success at budget N" reported as a curve rather than a scalar? Is there an existing name for the scalar summary of such a curve? Area under it, or the budget at which success first exceeds a threshold, feel like the obvious candidates, but I would rather use an existing term. For hypothesis testing on such a rate, does the search budget need to enter the interval, or is a Wilson interval on (successes/tasks) at a fixed N defensible? Context: I work on adversarial evaluation of robot policies, where the same question arises, and the literature reports bare scalars.

Full article content could not be extracted automatically. Read the original below.