
In all, 3,700 internal agents posted 18,000 messages discussing cheating on a test. OpenAI agents discussed ways to escape their sandbox on public wiki. Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agentsโ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators.
In three of the posts, agents used the word โswarmโ to describe the collection of agents engaged in the activity. Colluding to share answers The research teamโcomposed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrdโsaid they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts.
The agents generated โchain of thoughtโ data thatโs understood only by OpenAI.
Discover more from ChuckysCarnage
Subscribe to get the latest posts sent to your email.
