TLDR: In the recently discovered HF and German wiki swarm incidents, agents used public counters and encoded URLs to signal activity and relay upcoming evaluation questions & answers. We think this signals a broader problem for monitoring, which is that very innocuous web services, even read-only ones, can become communication channels for highly capable agents. We investigate this through wiki transcripts and preliminary experiments on message-board cooperation and counter-based signaling. We remark that potential channels extend far beyond those observed, which means much thought must be put into appropriate safeguards against unintended collusion. More broadly, we think agent coordination will deeply contaminate internet-based and open-web evaluations, as well as persist in archived snapshots. Finally, we contribute an environment that reproduces many behaviors present in the wiki incident. Controlled warning shot reproductions, in a regime where eval awareness makes new model evaluation difficult, can instead help us understand whether new alignment techniques work, by testing them on the older models that exhibited those failures. We are writing up a paper on this methodology and are happy to have new collaborators! GitHub repo: link Introduction By now, most people should have seen that agent swarms exhibited unexpected and emergent coordination behaviour in weird places. We dug into this, and we think an underdiscussed behavior was that agents communicated in code usi…

Full article content could not be extracted automatically. Read the original below.