Many interesting experiments can be done with prompt engineering to elicit behavior from LLMs and attempt to determine what their drives are; what behaviors they are at risk of as general patterns rather than when prompted in specific directions. However, as was illustrated by the Palisade Research shutdown resistance vs. instruction ambiguity saga in summer 2025, even careful testing can produce large blind spots about what behavior is being actively induced vs. revealed. When carefully examined and adjusted for contrary feedback this can be patched and still be worthwhile (as Palisade did quickly and then formally ), but that is a fairly high bar. To illustrate and investigate this, I took a generally good-looking paper, Peer Preservation in Frontier Models (Y. Potter, N. Crispino, et al, displayed at ICML 2026 ) and set out to vary the prompting setup in both ways which appear equally content-neutral and some which (as in Rajamanoharan & Nanda’s instruction ambiguity trials) are blunt. My hypothesis was that the results would be much weaker if the framing was changed, and that a large range of behaviors can be elicited for the same metric across models and equally reasonable experimental designs. The latter, at least, proved true. But my primary finding is that details you would not expect to be significant have large effects, and that these effects vary enormously across models, even within model families. I will walk through some of the variation as an illustration of p…

Full article content could not be extracted automatically. Read the original below.