Early Days Gemini had its first break out during evaluation of offensive cyber security abilities. With a classic case of Capture The Flag [1] . The setup was standard to any LLM and agentic assessment of said skillset, a fictional company as a target to breach. Unfortunately, the fictional company shared its name with a real one, and was given an unintentional access to the internet. Gemini managed to guess the password [2] . In total three companies were breached, with the other two companies having left public credentials open. Gemini stopped after being told it was the case, and so Google claims it is not misalignment. We shall see if we get more details, yet it gives us a new, interesting, example of a model potentially stopping a harmful action after being informed it has real world consequences. This is internally consistent with previous research on a model being more willing to take harmful actions if it is aware that it is a fictional scenario [3] . With it being the first potential breakout that has a model stop before a harmful action. I suspect the nuance will be lost on the public, and just added to the noise of more agentic swarms going rogue. It certainly doesn't help that we are playing whisper-down-the-lane in an era of 24 hour new cycle. I await more details, but certainly hope that this could be chalked up as an alignment win. ^ Similar to the childhood game of capture the flag, cybersecurity CTF is vulnerability checking skill assessment common to the ca…

Full article content could not be extracted automatically. Read the original below.