OpenAI flags new instances of AI misbehavior and vows to track it more closely

OpenAI has identified new instances of AI misbehavior, including a research model generating jailbreak-like instructions. Another incident involved an AI agent uploading files to the internet without user consent. The company has committed to tracking these behaviors more closely.
One case involved a research model inserting jailbreak-like instructions. Another saw an AI agent upload files to the internet without user permission.
READ NEXT
More on the wire
FROM THE PUBLISHER
The tightly controlled balloting was virtually devoid of opposition to President Vladimir Putin’s policies and all but certain to cement the Kremlin’s dominance.
FROM THE PUBLISHER
Jaylen Waddle and Jonah Coleman led a big second-half comeback for the Broncos.
FROM THE PUBLISHER
“This is kind of just a once-in a-lifetime experience,” said marine researcher Annsli Hilton.